Set your phasers to ambition
AI can raise your ambitions for pace, scope, correctness, reliability, and craftsmanship. Leaders get to choose where to spend the tokens.

The natural inclination, when talking about the benefits of AI, is to focus on speed. Execution of certain types of tasks, most notably writing code, has indisputably gotten faster. Things that used to take hours now take seconds. Things that used to take days now take minutes. Things that used to take weeks now take hours. We measure this in output: more code, more PRs, more features.
And it’s not actually just an AI thing. This has been the language, the driving ethos, of most software delivery teams for years. How do we get more software, to more customers, faster?
We prefer to talk about the benefits of AI in terms of ambition, and ambition can take many forms:
- Pace: How quickly you can deliver. Do the work you already planned to do, but faster (ideally on timelines that would have previously seemed outrageous).
- Scope: The scale or difficulty of what you attempt. Take on work that didn’t seem possible before.
- Correctness: Whether it behaves as intended. Amplitude reduced bug reports by 55% while tripling PRs.
- Reliability: How consistently people can depend on it. Intercom reduced downtime from breaking code changes by 35%.
- Craftsmanship: How thoughtful and refined the result is. If the feature can be prompted into existence over the course of a few hours (or even minutes), do you ship it as is or do you spend the time to refine it further?
These are not, of course, completely independent. They are frequently intertwined. Increased scope is only palatable because pace has dramatically increased as well. Craftsmanship is a type of expanded scope, and carries an expectation of correctness along with it.
The point is that you, as leaders, as managers, as ICs, get to decide where to spend your tokens. Spending them all on speed, on raw execution, isn’t your only choice, and it’s almost never the best choice for your business. Choose speed sometimes, sure. But also choose quality. Choose big bets. Choose ambition.
Here’s what we’re reading this week:
Tools Are Getting Better. Are We?
Karri Saarinen, co-founder and CEO of Linear, on product work, learning, and craftsmanship. A reflection on what we choose to do with the efficiency gains of AI, and what we stand to lose.
The answer is not to resist AI. My hope is with all the efficiency gains we are getting from the agents, we could actually spend more time to learn about our customers, problems and products.
Using the agents to understand more, not only produce more. Be deliberate about where and when does the learning happen for your team.
Write_On
If you care about product, do yourself a favor and watch this video from Jason Fried, co-founder and CEO of 37signals. He shows off a writing tool he built over the weekend. He’s not trying to sell it, he just built it to match the way he wants to work. But there are so many thoughtful features. It’s a great example of using AI to inflect craftsmanship and scope rather than just speed.
Vicent Martí on agents monitoring production deploys
“Last mile” agents seem to be the new hot trend in the agentic enterprise software loop: agents that “hand-hold” your change as it makes its way to production. We talked about this recently from inside OpenAI and Cursor recently released Rollouts to productize this capability. SpaceXAI engineer Vicent Martí doesn’t mince words about how transformative he thinks this use case is:
You can’t monitor a production deploy better than an agent can. I dunno about AGI but this specific job is something that an LLM can already do better than every single human in the world. For me, as a systems engineer, having agents monitor my infra changes has been (no joke) more impactful in my day-to-day work than having agents write my code.
Automating eval design and hillclimbing with Claude
Useful whether you use Claude or not. Good technical breakdown of what evals are, principles for eval design, and how to improve your skill or application by “hillclimbing” against an eval. Also details how to use the built-in /claude-api skills to put these concepts into action on your repo.
Is Jev overhyped? We tested it on 4 real enterprise tasks.
Well, we almost made it through the whole newsletter without talking about Jev. Sorry, but we’re still in the honeymoon phase. Don’t take our word for it, though. The Glean team gives a level-headed analysis of what Jev is (and isn’t) good for.
One thing to try this week
Agents love to write tests. This is, on net, a good thing. It also means in any codebase with active agentic contributions, you probably have more tests than you actually need. The OpenClaw team released a skill to audit your tests and get rid of “junk”. Took ten minutes to run against one of our repos and put up a PR to get rid of a handful of low-value tests.
Remember, nobody wins when you fight on Twitter.
— Claire + Zach