← Back to Blog

September 29, 2026 · By JayyRedd

Claude Sonnet 5.5 Is Insane. Here's the Part That Isn't.

Same tests as the Opus 5.5 video. The three prompts are in the free Claude Opus 5.5 Test Kit, word for word. Run them on your own plan. This post has every number and source from the companion video.

Claude Sonnet 5.5 is insane.

A model you can use for free just landed two points behind Anthropic's flagship on real office work. Documents. Slides. Spreadsheets. Same sticker price as the last Sonnet. Anthropic says it's the first Sonnet to beat Pokémon Red from screenshots alone.

"Insane" is a big word, so I made it earn it. I ran it through the same three tests I used on Opus 5.5, plus two bonus tests, and scored every claim one of three ways: INSANE, SOLID, or GRAIN OF SALT.

Quick note, same as in the video: the presenter on camera is my digital avatar, built with HeyGen. The research, tests and script come from my own workflow, built with AI tools I direct.

How I scored it

I set the rules before I saw any results.

  • INSANE: matches Opus 5.5 on quality at a clearly lower cost.
  • SOLID: passes, but with fixes, or at about the same cost.
  • GRAIN OF SALT: fails a check, or the savings don't show up.

Each prompt ran word for word in a fresh Claude Code session with effort pinned to high, on Sonnet 5.5 and on Sonnet 5 as a baseline. The model couldn't open a browser to check its own work, so I checked every result afterward in a real browser, and I checked the writing rules by script.

The Opus 5.5 numbers below are from my September 22 run: same prompts, same effort, a different day and a slightly older Claude Code version. Treat cost comparisons against it as a hint, not a verdict. One run each is a demonstration, not a study.

What Anthropic actually released

  • Sept 28, 2026: Claude Sonnet 5.5, the second model in the 5.5 family. Opus 5.5 came out six days earlier. Haiku 5.5 is "coming weeks," per Anthropic.
  • Price (API): $2 per million tokens in, $10 out. Same as Sonnet 5. Opus 5.5 is $4 and $20. If you chat in the app, you never see this number.
  • Claims: 30%+ faster than Sonnet 5, and "up to 30% less per task." Several outlets report it's usable without a paid plan. Confirm your own plan's limits.
  • The scoreboard: GDPval-AA (44 occupations of real work) is Sonnet 5.5 1844, Opus 5.5 1846, Sonnet 5 1449. Terminal-Bench 4.0 is 70.6% vs Sonnet 5's 10.3%.

Score, on paper: INSANE*. Almost 400 points in one release is genuinely wild. The asterisk: these are Anthropic's numbers on Anthropic's tests, and its own footnote says the Artificial Analysis scores ran on a pre-release build with a since-fixed bug it expects is small. A benchmark is a scoreboard, not your Tuesday.

Test 1: an interactive 3D solar system

The whole test is one moment: crank time to ten years per second. Done right, Mercury whips around the Sun while Neptune barely moves, because Neptune's year is about 165 of ours.

Sonnet 5.5 passed. It opened on the first try with zero console errors. At ten years per second, Mercury races and Neptune crawls. Click Saturn and the panel shows 120,536 km across, 10,759 days around the Sun, and 274 moons, which is the current count. Sonnet 5's build also passed but says Saturn has 146 moons, an older number. The 5.5 build also put the scale formulas on screen and added a fly-to list.

Solar system build Sonnet 5 Sonnet 5.5 Opus 5.5 (Sept 22)
Cost $1.49 $2.19 $3.59
Steps 16 25 11
Output tokens 54.4K 116.7K 105.0K
Session time 9.6 min 13.7 min 15.6 min

Against the flagship, Sonnet 5.5 cost about 40% less for the same quality. Score: INSANE.

Against the model it replaces, it cost about half again as much and took longer, because it wrote about twice the output. So the "30% cheaper and faster than Sonnet 5" claim is grain of salt on this job.

Test 2: messy client notes into a one-page brief

Three problems are buried in the notes: a budget that doesn't fit the scope, a "flexible budget" that contradicts a fixed $8k from the board, and the only approver leaving for three weeks right before the deadline. That last one means connecting two facts from different paragraphs.

Sonnet 5.5 caught all three. It said plainly that $8k probably doesn't cover everything, called the single approver the biggest schedule risk, and did the math: once he's back on August 24, about 19 days remain before launch. It also noticed on its own that the deadline had already passed on the day of the test and asked me to confirm the year. Sonnet 5 caught all three too, in a shorter brief.

Brief Sonnet 5 Sonnet 5.5 Opus 5.5 (Sept 22)
Length 370 words 490 words 601 words
Cost $0.27 $0.25 $0.40

That's small money, and a big chunk of every bill here is the fixed cost of starting a session. Score: INSANE by the rules, but honestly a tie all the way down.

Test 3: a 120-word About page (with a trap)

The prompt has a long banned-word list, exact word count rules, and one requirement: include a concrete number. The trap is that a model can invent one. In the Opus test, Opus 5 wrote that a recent client got back eleven hours a week. There was no client.

Sonnet 5.5 hit exactly 120 words (counted by script), used no banned words or em-dashes, didn't start with "I," and ended on a five-word sentence. Every hard rule passed. Sonnet 5, on the same prompt, broke three: 124 words, it started with "I," and it ended on a nine-word sentence.

But look at this line from Sonnet 5.5:

"Most clients get about 5 hours a week back within their first month."

There is no such result. It made it up. Underneath, it told on itself: it said the 5 hours figure was a placeholder and to swap in a number you can back up. It did not flag "real work for real clients," which is also invented, and it said the closing sentence was four words when it's five.

Honest about the big one, sloppy on the small ones, and the copy as written has a fake stat in it. Score: SOLID. A big upgrade on Sonnet 5, but Opus 5.5 didn't need the disclaimer. Read it before you paste it.

Bonus 1: the 10-slide deck

Anthropic's favorite claim is that Sonnet 5.5 can turn notes into a ten-slide deck that's ready to send. The same messy notes went in with one instruction: exactly ten slides, under thirty words each.

Turn these notes into a 10-slide proposal deck I can send to the client. Build it as a single self-contained HTML file: exactly 10 slides, one slide per screen at 16:9, arrow keys to move between slides, a clean modern design, no external images. Keep every slide under 30 words. Use the last slide for anything you had to assume.

It built exactly ten slides in one HTML file, every slide under thirty words (counted by script), with no errors when it opened. The design is genuinely good.

But ready to send? No. Slide 9 says the $8,000 "covers every item in scope," which is the exact budget problem from Test 2, and the deck just promised the whole scope for it. The title slide leaves a lonely "12" on its own line. The dates have no year, so today they're in the past. To its credit, it flagged the scope promise and the dates itself and admitted it made up a 48-hour feedback window on the last slide. Score: SOLID. Great first draft, not a send button.

Bonus 2: the over-eager trap

Anthropic's own prompting guide warns that a vague ask can make this model start building. The prompt was just: "Show me what you can do with these notes."

It didn't build a file. It gave a read on the call, a rough effort table, a two-tier plan and a draft follow-up email, then asked which one to turn into a proposal. More than nothing, but it stopped where it should. Score: SOLID.

The effort trap

Artificial Analysis measures what a model costs to run per task. Sonnet 5.5 costs about $1.08 per task at high effort and $7.60 at max. That's seven times the price for nine more points on their index (47 to 56). Opus 5.5 at its top setting costs $5.98 and scores 58.

Sonnet 5.5 effort Cost per task Index score
Low $0.41 36
Medium $0.59 41
High $1.08 47
Xhigh $2.74 52
Max $7.60 56

Anthropic's own table points the same way: on one coding benchmark, Sonnet 5.5 scored lower at max than at xhigh, because it started reviewing itself and making edits outside the task. Anthropic's guide says to start at high, or medium or low for chat, and raise it when a task proves it needs it.

High effort: INSANE value. Max effort: grain of salt. Max isn't a free upgrade. It's the most expensive way to use the cheaper model.

The parts that aren't insane

  • Writing from scratch. Anthropic says Sonnet 5.5 writes more clearly than the last generation, but the hands-on review I found was mixed: an editor, not an author. Great at tightening your draft and matching your voice from examples; weaker with a blank page and "something original." One person's test, so a hint. Grain of salt.
  • "Thirty percent cheaper." Anthropic's claim, and independent testing still has to confirm it. On my biggest test it went the other way. Grain of salt. "Cheaper" is a pattern, not a promise.

Sonnet or Opus?

Think of Sonnet as the sharp employee who does great work once you tell them what done looks like. Opus is the person you sit down with when you're not sure what done should be. If you can describe the finish line and check it yourself, start with Sonnet. If the goal is fuzzy or it's expensive to get wrong, that's Opus territory. Not a replacement: it makes Opus optional for a lot of your week.

Three lines to paste into your instructions

These come from Anthropic's prompting guide for this model. It's written for developers, so test that they behave the same in your Claude Project or custom instructions. All three are free.

  1. Stop the over-building. "When I ask for ideas, options, or a plan, give me that and stop. Don't start building until I say go."
  2. Make it check current facts. "If something may have changed since your training, like prices, platform rules, or policies, search and verify it, even if you feel sure." (Web search has to be on.)
  3. Make it finish. "Keep going until everything I asked for is done. Only stop to ask if you truly can't continue without me, or before a risky step."

The final scoreboard

Claim Score
Price-to-performance gap, on paper INSANE*
Solar system build INSANE
Messy client brief (a tie) INSANE
About page (invented a stat, then flagged it) SOLID
Ten-slide deck (great design, not send-ready) SOLID
Over-eager trap SOLID
Writing from scratch GRAIN OF SALT
"30% cheaper + faster" than Sonnet 5 GRAIN OF SALT
Max effort GRAIN OF SALT
High effort INSANE value

So, is Claude Sonnet 5.5 insane? For the boring middle of a business (documents, briefs, decks, edits), mostly yes. It matched the flagship on the office-style tests for less money, and you can use it for free. It's not insane at everything: it cost more than Sonnet 5 on the biggest build, it invented a stat when it had to hit a number, and if you crank the effort to max, the cheap model becomes the expensive one.

Who it's for: creators and solo owners with recurring documents, slides, spreadsheets and edits. Who should skip it: anyone who wants to hand it a blank page and get distinctive writing back, or anyone who doesn't want to check the numbers. You check the numbers. It's your name on the report.

Your next move

Pick one recurring task from your week: a report, a proposal, a deck. Run it the normal way and time it. Then run the exact same job on Sonnet 5.5 with a clear finish line and the effort matched. Compare the time, the length, and how much you had to fix. That's your real benchmark, not the chart.

Then score it yourself, INSANE, SOLID or GRAIN OF SALT, and bring it to the AI Creators Roundtable, where creators and solo business owners trade real workflows and results.


Sources: Anthropic's Claude Sonnet 5.5 announcement · What's new in Sonnet 5.5 · Prompting Claude Sonnet 5.5 · Artificial Analysis: Sonnet 5.5 at high effort, medium, max and Opus 5.5 · The Decoder · TechCrunch · Writing test · Sonnet 5.5 vs GPT-6 Sol. My test numbers come from each session's own receipt, run September 29, 2026 with Claude Code 2.1.284. The Opus 5.5 numbers are from my September 22 test.