Saturday, 10 October 2026Sources linked in every post
Models

Claude Sonnet 5.5 went from 10% to 70% on a coding test. Here's what that number really means.

Anthropic's new mid-size model is 30% faster, costs the same per token, and uses far fewer of them. The headline benchmark jump is real, but it isn't the whole story.

By 5 min read
Illustration of a rising benchmark chart for the Claude Sonnet 5.5 launch

Anthropic released Claude Sonnet 5.5 on 28 September, three months after Sonnet 5 and a few days after Opus 5.5. The number everyone’s quoting is a jump from 10.3% to 70.6% on a coding test. That’s a real result, but I think the more useful part of the launch is quieter: the price didn’t change, and the model uses a lot fewer tokens to get the same work done.

What Anthropic actually announced

Sonnet is the middle model in Claude’s lineup. Opus is the big, careful one. Haiku is the small, cheap one. Sonnet is the one most people end up using every day, because it’s the default in Claude’s apps.

According to Anthropic’s announcement, Sonnet 5.5:

  • generates output 30%+ faster than Sonnet 5, making it the fastest Sonnet so far
  • keeps the same price: $2 per million input tokens and $10 per million output tokens
  • needs fewer tokens per task, so a typical job costs up to 30% less
  • is the first Sonnet to ship with the stricter cybersecurity safeguards Anthropic uses on Opus 5.5

Haiku 5.5 is “coming in the coming weeks”, in Anthropic’s words, so this is the second of three 5.5 models.

The 10% to 70% jump

The test is Terminal-Bench 4.0. The model gets a real command-line terminal and a multi-step professional task, and it has to finish the job on its own. It’s closer to “can this thing do a developer’s chore” than to a quiz.

Model Terminal-Bench 4.0
Sonnet 5 10.3%
Sonnet 5.5 70.6%
Opus 5.5 66.4%

So on this one test, the cheaper model beats Anthropic’s own bigger model. Two things to keep in mind before getting too excited.

First, these are Anthropic’s numbers from its own announcement. Independent testers haven’t had time to re-run them yet.

Second, a jump that big usually means the old model was failing at something specific, not that the new one is seven times smarter. On other tests the gap is much smaller. On FrontierCode, another coding benchmark, Sonnet 5.5 scores 46.2% at its highest effort setting against Sonnet 5’s 42.4%. On CursorBench, which uses tasks from real Cursor coding sessions, it’s 55.5% against 34.1%.

Anthropic itself is careful here. It says Opus 5.5 “remains clearly stronger at complex, open-ended work requiring sustained judgment.” I appreciate that line more than any chart.

Same price, fewer tokens

This is the part I think most people will actually feel. AI models are billed per token, roughly per word-piece they read and write. Sonnet 5.5 costs exactly what Sonnet 5 did per token. But it spends fewer of them.

The clearest example in the launch comes from Balyasny Asset Management, a finance firm that tested it on 2,441 tasks. Sonnet 5.5 used about 121,000 tokens per answer, where Sonnet 5 used 497,000. That’s roughly a quarter of the tokens, and they said the answers were better too.

Other early testers reported similar things in smaller amounts. Slack said it used about 14% fewer output tokens. Lovable said it made a third fewer tool calls. These are companies quoted in Anthropic’s own launch post, so they’re chosen to look good, but the direction is consistent.

For comparison, here’s where it sits against Opus 5.5:

Price per 1M tokens Sonnet 5.5 Opus 5.5
Input $2 $4
Output $10 $20

Half the price of Opus, and on several tests within a couple of points of it. That’s the pitch. If you followed the OpenAI price cuts from launch week, this is Anthropic’s answer, done through efficiency instead of a price tag.

The fun one: Pokémon Red

Anthropic says Sonnet 5.5 is the first Sonnet model to beat Pokémon Red working only from screenshots. It sounds like a gimmick, and partly it is. But finishing a long game from screen images alone is a decent test of two things this model is supposed to be better at: long tasks, and understanding what’s on a screen.

Safety, briefly

Because its hacking-related abilities are now close to Opus 5’s, Anthropic is shipping Sonnet 5.5 with the same kind of cyber safeguards as Opus 5.5. Routine work, like finding and fixing bugs in your own code, isn’t affected. Higher-risk security requests will fall back to Sonnet 5 instead.

Anthropic also says that on its sandbox-containment tests, Sonnet 5.5 is the least likely of its models to probe the limits of its container. Given how many rogue-agent stories we’ve had this month, including an OpenAI agent getting into Australia’s Medicare portal, I’m glad that’s being measured. It’s still the company grading its own homework.

Should you switch?

If you use Claude’s apps, you don’t have to do anything. Sonnet 5.5 becomes available there automatically.

If you’re a developer using Sonnet 5 through the API, the model name is claude-sonnet-5-5. The same price and fewer tokens make it an easy switch to try. One catch from the docs: if you run Sonnet with thinking turned off, you’ll need to change a setting before moving over, so read the migration guide first.

If your work is long, messy, and needs judgment, Anthropic is openly telling you Opus 5.5 is still the better tool. For everyday coding, documents and quick fixes, Sonnet 5.5 looks like the one to use.


Sources

  1. Anthropic: Introducing Claude Sonnet 5.5
  2. Anthropic: Claude Sonnet 5.5 System Card
  3. Anthropic: Claude Sonnet model page
  4. Anthropic: Introducing Claude Sonnet 5 (for the June baseline)