TheoSym

Claude models compared

Claude Opus vs Sonnet: which model should you use?

Make Claude Sonnet 5.5 your default and move up to Opus 5.5 only when the task is complex enough to need it. In the Terminal-Bench 4.0 numbers we covered, Sonnet 5.5 scored 70.6% against 66.4% for Opus 5.5 at about half the price per token, while Opus still leads on FrontierCode and CursorBench.

Updated September 30, 2026 · By Dr. Sam Sammane

From the TheoSym channel

Sonnet 5.5: the new era of cheap Claude

Published September 29, 2026

Sonnet 5.5 at low or medium effort beats Sonnet 5 at its best for about a tenth of the cost per task. More effort means longer thinking and a bigger bill, so start low and climb only when the work needs it.

Watch on YouTube

Claude Sonnet 5.5 vs Opus 5.5 at a glance

Sonnet 5.5Opus 5.5
Input price (per 1M tokens)$2$4
Output price (per 1M tokens)$10$20
Terminal-Bench 4.070.6%66.4%
Where it leadsAgent coding tasks on Terminal-Bench 4.0, at half the priceFrontierCode, CursorBench and complex work

Prices and scores are as reported in our videos of September 29–30, 2026, citing Anthropic. Confirm current prices on Anthropic's pricing page before you budget.

How to choose between Sonnet and Opus

The lesson from the numbers is to pay for results, not for the name. A cheaper model that finishes the job is a better buy than a more expensive one that finishes it slightly better.

  • Start with Sonnet 5.5 for day-to-day coding, agent loops, drafting and analysis.
  • Move to Opus 5.5 when a task is complex enough that Sonnet fails or needs several retries, or when it matches an area where Opus still leads, such as FrontierCode or CursorBench.
  • Decide per task, not per team. Route routine work to Sonnet and reserve Opus for the hard cases.

Effort level changes your bill more than the model name

In our video on Sonnet 5.5, the point is that Sonnet 5.5 at low or medium effort beats Sonnet 5's best score for about a tenth of the cost per task. More effort means longer thinking and a bigger bill.

So the practical rule is to start at low effort and climb only when the work needs it. Test the same task at low, medium and high before you set a default.

The cost you do not see: thinking tokens

You pay for every token a model writes, including the reasoning you never read. In the finance benchmark we cited, Balyasny ran 2,441 tasks. Sonnet 5 used about 497,000 tokens per answer, and Sonnet 5.5 used roughly a quarter of that.

When you compare models, compare cost per finished task, not just price per token. A model that thinks less can be cheaper overall even at the same rate.

Test on your own work before you switch

Benchmarks measure someone else's tasks. Anthropic itself says Sonnet 5.5 "may have tendencies we haven't found," so treat published scores as a starting point.

  1. Pick 20 real tasks from your own work.
  2. Run each on Sonnet 5.5 and Opus 5.5 at the same effort level.
  3. Record whether each result was usable, how many retries it took and what it cost.
  4. Set Sonnet as the default and route to Opus only where it clearly earns the extra cost.

Use our free LLM API cost calculator and model migration cost estimator to put your own numbers on the switch.

Claude Opus vs Sonnet: common questions

Is Claude Sonnet 5.5 better than Opus 5.5?

On Terminal-Bench 4.0, using the numbers we covered, Sonnet 5.5 scored 70.6% against 66.4% for Opus 5.5 on agent coding tasks. Opus 5.5 still leads on FrontierCode and CursorBench, so the better model depends on the task.

How much do Claude Sonnet 5.5 and Opus 5.5 cost?

As reported in our videos, Sonnet 5.5 is $2 per million input tokens and $10 per million output tokens. Opus 5.5 is $4 in and $20 out. Check Anthropic's pricing page for current rates.

Which Claude model is best for coding?

Start with Sonnet 5.5 for most coding and agent work, since it scored higher than Opus 5.5 on Terminal-Bench 4.0 at half the price. Move to Opus 5.5 for the hardest tasks where Sonnet needs many retries.

What does the effort setting do?

It controls how long the model thinks. Higher effort means more thinking tokens and a bigger bill. Sonnet 5.5 at low or medium effort beat Sonnet 5's best score for about a tenth of the cost per task, so start low.

Should I switch my agents from Opus to Sonnet?

Test first. Run your own tasks on both models at the same effort level, compare cost per finished task, and switch only where Sonnet gives usable results.

Want an agent you can trust in production?

TheoSym ships production agents with the eval suite, MCP tools and harness included. Bring one real workflow to a 15-minute call with Sam and see what building it would involve.

Book 15 minutes with Sam