← Notes
Weekly notes · · Claire Vo + Zach Davis

Glimpses of the future

Early agent-native workflows reveal that durable gains come from repairing the friction between agents, tools, context, and human review.

Pixel-art robot in a dark control room tagging colored nodes on a glowing wall-sized network map, alien city through the window

You don't need us to tell you that the future is uncertain. That's both the thrill and the torture of the moment we're in. And if you're close to this juggernaut we're calling AI (and, truly, why would you be reading this if you weren't) you feel one or both in your bones every day.

Some people enjoy prognosticating more than others (even amongst the two of us; if you know us, it's not hard to guess which is which). But when you spend as much time as we do both using these new classes of tools for real and also talking to companies in various stages of AI transformation, you really do start to see glimpses of the future before they become widely accepted. Often these don't feel like predictions; it's more like you're just naming what you've experienced or what you see, before the mainstream has fully caught up.

And so it's with some trepidation that we're going to suggest you read at least one of Steve Yegge's latest lengthy reports from the forefront of AI psychosis. But also, dude is smart and whether or not his conclusions are correct, he's operating at the bleeding edge and letting us know what he sees from there.

Will CI/CD go away? We wouldn't bet on it. In fact we see it as a great complement to agentic verification. But verification in general is ripe for disruption and will probably look very different over the next year.

Will human review go away? Probably not entirely, but what seemed unthinkable a year ago (serious enterprise companies regularly merging a subset of changes to production without any human review) is now an emerging standard practice that will only continue to grow in both kind and breadth.

Is Yegge's "Wish Factory", which is really not that different from the dark factory that so many are salivating after, an inevitability? For a certain class of changes, yes. If you can ship bug fixes, small improvements, and other well-vetted changes to customers with minimal friction, including the friction inherent in human involvement, that's a competitive advantage (or perhaps soon it will be a competitive disadvantage if you don't). Will everyone turn their roadmaps over to their customers? Almost certainly not.

Does anyone have all the answers? Of course not. Are there people out there who know more than you, or have experience with things you don't (yet)? Yeah. Should you listen to all of them? No. But if you're not prioritizing learning and experimenting, it is not hyperbole to say that you're going to get left behind. Or to put a more positive spin on it: we're in a moment where trying something very few people have ever done before is easier than ever, and all it requires is you being willing to meet that moment head on.

Here's what we've been reading this week.


The Shape of Things to Come

Steve Yegge reporting from his enclave in the future. He says the secret to loops and graphs is infinite tokens (it definitely helps). As discussed above he thinks CI/CD is dead (just merge to main and sort it out there) and human review is finished (certainly diminished). It's a lot to take in but it certainly makes you think.

That's the shape of things to come. It ain't gonna be a framework you download, or a harness from someone who's not building an actual thing. You're going to be building a whole civilization, plank by plank, with colleagues who happen to run on datacenter silicon.

If you read all 6,700 words and think to yourself "I'd like more of that, but even more out there", well, you're in luck. He also published part two: Model Welfare for Agentic Engineers

What I Want to Tell You About Orbs

Thorsten Ball, who is always a delight to read, extolling the virtues of the new Amp feature "Orbs". Orbs are, as far as we can tell, just the infrastructure that allows Amp to support background agents. They may be extremely well crafted background agents, but they appear to just be background agents. The Amp team is collectively losing their mind on social media about them, including in the post linked here, which admittedly makes us curious.

Here’s what you get with orbs: more agents building more complicated things; agents running for longer and giving better proof that what they did works; a lighter, less sigh-inducing review load; more things shipped, faster.

>

Now I just need to figure out how to make you believe me.

But also we are on-the-record long-running fan-kids of Cognition's Devin, the original (and occasionally maligned) background agent, so mostly we're just happy that the Amp team is validating what we've been telling people for over a year: background agents are a major unlock in the world of agentic engineering and well worth the investment.

Teaching an LLM to review code like a Senior Engineer!

Intercom's Kesha Mykhailov gave a talk at WeAreDevelopers World Congress 2026, providing ample insight into their code review and auto-approval bot Shrek. We've talked a fair bit about Intercom's approach here, but this talk gives more detail than we've seen so far in the wild.

Kesha gives lots of technical details, including some wrong turns and some hard-won lessons (evals, backtests, and a/b tests). But he also talks a lot about the "socio-technical system", or the incentives that get emphasized or reinforced through the tooling. This is easy to overlook, but it's the difference between designing a technically sound system that nobody likes and designing a system that actually gets used.

He talks both about the first-order and second-order effects, but the first-order socio-technical effect is our favorite: no pull request over 150 lines of code can get auto-approved, directly incentivizing a well-known but hard to enforce rule of thumb that has plagued engineering orgs for decades: small changes are easier to reason about, safer to merge, and just generally better engineering practice.

Build an AI code review agent with Vercel Eve

Claire's How I AI episode this week also tackled a custom review and approval agent. It's a good complement to Kesha's talk. Where Intercom spent months building and tuning their review agent earlier this year for a large enterprise rollout, Claire shows how state-of-the-art tooling and know-how allowed her to build an agent in just a few hours that could effectively review and "approve" PRs for ChatPRD.

And Claire herself says that several months ago she tried the same thing and ultimately felt it wasn't worth the effort. Now? A morning gets you something real you can play with, so you can spend the afternoon getting your agent to apply Intercom's lessons and writing evals and backtests.


One thing to try this week

Automated friction logging for agents. Agents have gotten much better at running for long periods of time. In practice this mostly means they've gotten more resilient to failure. When they run into a problem they're more likely to recover, try something else, and move past it. This is great if you want an agent to finish a task while you're not paying attention. It's less great if you want to stop the friction at the source so all your agents don't run into the same problems over and over.

Steve Ruiz of tldraw shared his low-tech solution while somebody else packaged up a similar idea into a GitHub repo called Frog. Pick one and give it a whirl to find out where your agents are struggling without you noticing.


Keep calm and embrace the fun.

— Claire + Zach

Get the next issue in your inbox at CXO.dev Notes.

Let's go.

If you're stuck, we want to help. Tell us what's working (and what's not) and we'll get you where you need to be.