Anthropic's Claude Sonnet 5.5 is a quicker, cheaper way to run its mid-sized model. It came out on Monday 28 September as the second model in the Claude 5.5 family. Anthropic says it runs more than 30% faster than Sonnet 5 and costs up to 30% less for most work.

The token prices haven't changed. Anthropic says the saving comes from the model using far fewer tokens to get the same job done. It's out now on the Claude Platform as claude-sonnet-5-5, and on Amazon Web Services, Google Cloud and Microsoft Azure.

Where Sonnet 5.5 fits

Sonnet sits in the middle of Anthropic's lineup. The company launched Claude Opus 5.5, the first 5.5 model, on 22 September. It said Opus 5.5 performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5. Fable 5.1, which Anthropic released on 1 September, scored 66 on Artificial Analysis's Intelligence Index, ahead of Opus 5.

Photo: Sai Kiran Anagani _imkiran / Wikimedia Commons (CC0)

Anthropic's launch post pitches Sonnet 5.5 as "a faster, lower-cost complement to Claude Opus 5.5." Anthropic says Opus is for complex work that needs careful judgment. Sonnet's best at clear, everyday tasks, fixing bugs, and turning out polished documents, slides and spreadsheets. Claude Haiku 5.5, built for huge volumes of work where cost matters most, is due in the coming weeks.

The benchmark numbers

These scores come from Anthropic's own table, and the biggest jump is in agentic coding, where the model works through coding jobs on its own. On Terminal-Bench 4.0, which tests multi-step professional work in a command line, Sonnet 5.5 scores 70.6%. Sonnet 5 scored 10.3%. Anthropic lists Opus 5.5 at 66.4% on the same test at its highest effort setting. On CursorBench 4.0, built from real Cursor coding sessions, Sonnet 5.5 scores 55.5%, up from 34.1%. Anthropic lists Opus 5.5 at 57.8% on the same test, so Sonnet's best score is within about two points of its bigger sibling.

Office-style knowledge work improved too. GDPval-AA tests real-world tasks across 44 occupations and nine major industries. There, Sonnet 5.5 scores 1844. That's almost level with Opus 5.5 on 1846, and about 400 points above Sonnet 5. On the OSWorld 2.1 test, which checks how well a model uses a computer, it reaches 80.1%, up from 57%. Anthropic also says it's the first Sonnet model to beat Pokémon Red working only from screenshots.

Anthropic gave an example from its own testing as well. It handed Sonnet 5.5 a public company's quarterly earnings materials and call transcripts, plus a slide template, and asked for a 10-slide operating review. The company says two experts judged its first attempt ready to send as it was.

Anthropic adds its own warning. In its testing and in outside testing, it says, "Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment."

For anyone paying the bills, the cost charts might matter more. Anthropic says that on several benchmarks, Sonnet 5.5 at Low or Medium effort beats Sonnet 5's best score for about a tenth of the cost per task. FrontierCode checks whether an agent's code changes would get merged. At High effort there, Anthropic says Sonnet 5.5 scores 10 points higher than Sonnet 5 on the same setting, at about one-fifteenth of the cost per task. Medium is the default effort in the Claude apps and Claude Code. High is the default on the Claude Platform.

What early testers said

Anthropic published comments from companies that tried it before launch. "Claude Sonnet 5.5 cooks," wrote Tyler Nishida, a designer at Every. He said it's fast at coding but can still work for a long stretch when it has to.

Slack principal engineer Curtis Allen said that without changing any prompts, Sonnet 5.5 beat Sonnet 5 on almost all of Slack's offline Slackbot evals, "in fewer steps and with about 14% fewer output tokens." Zendesk's director of AI, Abhinay Kathuria, said it made fewer wrong calls on hundreds of real support cases. "Tickets were processed 20% faster," he said.

Epic Games chief operating officer Daniel Vogel said that in Epic's early testing it "cleared the same quality bar you'd expect from a higher-tier model." Unity creative technologist Sam Zhang said it finished 90% of tasks in Unity's multi-step Editor and coding benchmark. Lovable co-founder and CTO Fabian Hedin said its coding evals showed "a third fewer tool calls and roughly half the shell runs to finish a task." Another tester, who isn't named in Anthropic's post, said that across 2,441 finance tasks Sonnet 5.5 used about 121,000 tokens per answer. Sonnet 5 used 497,000.

Atlassian's head of product for AI, Jamil Valliani, said teams will be able to run Rovo agents up to 30% faster than with Sonnet 5.

Safety, safeguards and price

Anthropic says Sonnet 5.5's cyber skills are on par with Opus 5's. That makes it the first Sonnet model to launch with cyber safeguards like the ones on Anthropic's most capable models. Everyday work like finding and fixing bugs in your own code isn't affected. But higher-risk cybersecurity tasks "will visibly fall back to Sonnet 5," the company says. Cyberdefenders will soon be able to apply to a bigger Cyber Verification Program for tiered access to more advanced features.

It's also the first Sonnet with safety classifiers that block reasoning extraction. That's when attackers use thousands of fake accounts to copy a model's abilities. It's the same kind of campaign OpenAI said this week it had shut down against its own models.

On alignment, whether the model does what people actually want, Anthropic's automated behaviour check runs Claude through roughly 1,850 scenarios. The company says Sonnet 5.5 matches or beats Sonnet 5 on most measures. It says it found "no evidence that Sonnet 5.5 pursues goals that conflict with the user's intention." Anthropic also says it's the least likely of its models to test the walls of the containers it runs in. And it notes that no set of tests catches every failure.

Pricing stays at US$2 per million input tokens, US$10 per million output tokens and 20 US cents per million cache reads. That's half what Opus 5.5 charges for input and output. It's a crowded price point this week. OpenAI's GPT-6.1 Sol, launched at DevDay a day later, carries the same US$2 and US$10 rates on OpenAI's API model page. And Google set the same introductory rates for Gemini 4 Argon on 30 September.

There's one practical point for developers. If you run Sonnet with thinking switched off, Anthropic says you'll need to move to a new between_tools setting before switching to Sonnet 5.5. That setting keeps up-front thinking off. Like Opus 5.5 and Sonnet 5, it's available with zero data retention.

Anthropic has put out a short official video, "Introducing Claude Sonnet 5.5", on its Claude YouTube channel.