← Notes
Weekly notes · · Claire Vo + Zach Davis

Cheaper models, bigger bets

Cheaper frontier models and Jev's decision model invite leaders to revisit ambitious AI experiments and rethink which work needs text generation.

A split seed pod with four winged seeds drifting to the right.

We usually leave model-release takes to "How I AI"; if that's what you're looking for, it's our wholly biased opinion that you can't do much better. Over here in this neck of the woods, in this newsletter, we're more concerned with things that change how companies work, and there's a natural lag between frontier model releases and the innovations and trends that those models unlock.

But we sure got some notable model releases in the last two weeks, and it sure seems relevant to AI transformation, and, well, color us intrigued for a few reasons. First, it seems like suddenly there's a fire sale on intelligence and that seems poised to break through the cost wall the industry has erected in recent months. And second, Jev Jev Jev Jev Jev Jev Jev. Oops, I meant to say: we have also been swept up in the Jev hype and we're pretty excited about the possibilities (especially for the enterprise)!

The great cheapening

When Anthropic launched Fable 5, after months of Mythos whisperings and conjecture, they talked primarily about "capabilities", saying:

It is state-of-the-art on nearly all tested benchmarks of AI capability, showing exceptional performance in software engineering, knowledge work, vision, scientific research, and many other areas.

When OpenAI launched GPT-6 Astra, "a new generation of intelligence", they touted it as "the world’s most intelligent and aligned model". This week we got releases from both, and the tune is no longer just about capability or intelligence. Anthropic said:

We’re introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.

OpenAI said:

That’s why we’re expanding the GPT‑6 universe with GPT‑6 Sol and GPT‑6 Luna. GPT‑6 Astra introduced a new generation of intelligence—these models help distribute the benefits of that intelligence by advancing the frontier on cost efficiency.

Suddenly, we're talking about price. We're getting frontier-level intelligence at a fraction of the cost. That matters because cost comes up in every single conversation we have with leaders trying to turn AI experiments into regular work. On Anthropic’s CursorBench test, Opus 5.5 at medium effort edged out Fable 5.1 at max effort, at a lower cost per task.

So what does this mean for you, leader at Acme Corp trying to balance next year's budget with investment in AI transformation? It means your most ambitious projects, the ones that seemed out of reach 3 months ago, are now in play and should be revisited. Can you afford a two week experiment with an open-ended Opus 5.5 or GPT-6 Sol budget? We think you can and should.

The great Jev-ening

Jev took the ~~world~~ internet by storm last week. If you want to learn more about Jev, we'll link to someone smarter and more eloquent than us below. For now all you need to know is that Jev is an AI model, but instead of long streams of surprisingly intelligible text it returns surprisingly effective decisions. Yes or no. Which group does this belong to. That kind of thing.

So the internet lost its mind because we collectively spent the last two years spending all of our time trying to figure out how to apply intelligent-ish text generation to every problem ever invented, and suddenly we had a new toy to try to solve all our problems with. Suddenly every problem didn't need to have an LLM-shaped solution, it could have an LLM-shaped solution or a "system one model"-shaped solution.

The thing is, though, that Jev isn't just a shiny new toy. It's a really useful shiny new toy. And it’s cheap. The interesting question isn’t “Can Jev replace an LLM?” It’s “Which decisions never needed a text-generating model in the first place?” and "How can we combine the power of an LLM with the speed and frugality of Jev?"

Ultimately, a lot of what we work with companies on is bottlenecked on people, not technology. The technology is here, and the cost has been slashed. It's worth asking yourself if you're the bottleneck.

Here’s what we’re reading (and watching) this week:


Claire on Opus 5.5 and Opus 5.5 vs GPT-6 Sol

If model takes are what you want, here's where to get them. Opus 5.5? Not so annoying. Great at frontend design and SVGs. But can it earn a place in Claire's heart, or will she stick to her beloved OpenAI models? She did a live benchmark to find out.

Jev introduces a new shape of LLM — System One, aka Decision Models

Simon Willison with a level-headed Jev explainer. Skip the hype, get the facts. Actually, the hype is pretty fun. Get the facts, then go get some hype (it's easy to find), then go try it for yourself.

I think the decision model framing is useful for understanding where to use Jev. It’s great for anything that can be expressed as a classification task—think spam detection, suggesting labels, prioritization and ranking.

Harbor: Stripe’s AI-assisted prototyping tool

Staff engineer Cristian Rivera and technical writer Sai Samant detail the motivation behind building a custom prototyping tool at Stripe. They started by making it easier to fully prototype changes to stripe.com, but saw it get picked up and used by finance, operations, and sales.

There's also some interesting technical details about how they handle rendering, how they anchor comments to rapidly changing UI, and integrating with other tools. So many companies with the right culture and right technical foundations are creating bespoke AI tooling to solve their needs in ways that are tightly coupled to their technologies and the ways they work.

How we made claude.ai 3x faster in two weeks

The Claude team shared how they dramatically improved the performance of both the web app and the desktop app in just two weeks. It's a fascinating peek into how Anthropic is using the advanced capabilities of the Claude harness, alongside SOTA models, to do ambitious things.

A few things stand out. First, they organized around a Slack channel using Claude Tag, which means they worked collaboratively, in public, using a cloud agent. All signs point to something like this being a big part of future work. Second, they didn't just tell Claude to improve performance, make no mistakes. It was a collaborative effort between humans and agents, and also amongst humans. As we noted recently, teams of humans and agents working together are more effective than humans or agents working alone.

Mike Cannon-Brookes on building a software factory one station at a time

There's a lot of talk about software factories right now. Frankly, most of it is bullshit.

Never one to mince words (an early Atlassian value was "Open company, no bullshit"), Atlassian CEO Mike Cannon-Brookes delivers a sharp rebuke to the factory hype while simultaneously providing encouraging words to anyone with factory-like ambitions. Start small, automate what you can, keep moving forward.

You already have a software factory. AI just allows you to make more steps repeatable, more flexible - and that lets you move where you stand.

One thing to try this week

Claire used Jev and Gemini Flash to do a data deep dive on ChatPRD PRs to surface investment by theme, for just 9 cents. Her take?

As a CTO I would have paid $100k for this 3 years ago

Have fun. Be ambitious.

— Claire + Zach

Get the next issue in your inbox at CXO.dev Notes.

Let's go.

If you're stuck, we want to help. Tell us what's working (and what's not) and we'll get you where you need to be.