Claude Opus 5 Is for the Work Nobody Double-Checks
I've watched a practice group spend forty minutes arguing over a six dollar difference. Meanwhile voice mode showed up on every associate's phone, and nobody at the firm decided that.
Claude Opus 5 Is Here. Pick Your Model by the Cost of Being Wrong.
TL;DR: Opus 5 shipped July 24. Same price as Opus 4.8, $5 and $25 per million tokens, and it tops Anthropic’s reported coding and knowledge work benchmarks. Voice mode got Opus and Sonnet the day before, along with reach into your email and calendar, which is the part your firm should actually be worried about. Sonnet 5 is enough for most of what you do. Opus 5 for the work nobody’s going to double-check.
The notification hit while I was halfway through a drive up to Paso Robles for a wine weekend, and my first reaction wasn’t excitement. It was a small, tired sigh.
If you run a firm, you know the feeling. “Keeping up” stopped being a reasonable goal sometime around March. Sonnet 5 landed June 30. Fable 5 and Mythos 5 arrived June 9, then Fable got pulled for nearly three weeks over export controls and came back July 1. Now Opus 5. Four significant releases in under seven weeks, every one of them arriving with a chart showing it beat the last one.
I’m going to skip the chart.
What Anthropic actually shipped
Opus 5 costs exactly what Opus 4.8 cost. Five dollars per million tokens in, twenty-five out. Same sticker, better model, which is not how this usually goes. It’s the default now on Claude Max and the strongest option on Claude Pro, and Anthropic’s claim is that it lands close to Fable 5’s intelligence at half Fable’s price.
The legal-adjacent numbers are the ones I’d actually read. A contract AI vendor testing first-turn redlines said Opus 5 scored highest of anything they’d tried, close to double Opus 4.8. Box measured 17 percent better than Opus 4.8 on due diligence workflows. A second legal vendor saw its biggest gains in corporate governance and arbitration, and matched Opus 4.8’s max-reasoning quality on 26 percent fewer tokens, which matters more than it sounds if you’re paying by the token. All customer testimonials from Anthropic’s own launch post, by the way. Nobody independent has run any of it.
Then there’s an alignment result nobody covered. Anthropic runs an automated behavioral audit on its own models, and Opus 5 came out lowest of any recent model on misaligned behavior. Lowest rates of deception. Hardest to talk into something it shouldn’t do. I’d trade five points of coding benchmark for that in a heartbeat, because the person handing this thing privileged material at 11pm is not going to be the person who understands what it is.
Everyone’s arguing about the wrong number
I ran the cost numbers because someone always asks.
Two million input tokens in a month, two hundred thousand out. That’s heavy drafting and review for a small firm. Opus 5 comes to roughly $15. Sonnet 5 at its standard rate, roughly $9.
Six dollars.
I have watched practice groups spend forty minutes on this. And most of you are on subscriptions anyway, buying usage limits rather than tokens, so the whole exercise is theater.
So price isn’t the variable. Difficulty isn’t either, though that’s the one people reach for second. Some of the worst AI failures I’ve watched land in firms came out of work nobody would call hard. Just a lot of it.
Pick by what it costs you to be wrong.
Opus 5 or Sonnet 5: a working rule
Sonnet 5 handles anything you’re going to read anyway. First drafts, summaries of your own notes, reformatting a deposition outline, research you’d verify no matter who produced it. It’s the default on Free and Pro, it’s quick, and at $2 per million input tokens through August 31 it’s cheap enough to leave running in the background.
Opus 5 is for work you can’t check. Not hard work. Unverifiable work.
That gap is wider than it looks. If an associate runs eight hundred documents for privilege, nobody is re-reading eight hundred documents. And if a model runs a twenty-step research task, an error in step three quietly poisons everything downstream while looking entirely reasonable on the page, which is the failure mode that actually scares me. The Friday 6pm redline going to opposing counsel with no second pass belongs in the same bucket. Six dollars is not a consideration for any of these.
My own first test on a new model isn’t serious. I ask it to plan a week of meals for a 165-pound St. Bernard with loves freshly BBQ’d steak. Magnus is not a benchmark. But I’ve run the same prompt on every model for two years now, and inside a paragraph I know whether it’s going to hedge itself into uselessness or commit to an answer. Opus 5 committed. Make of that what you will.
Voice mode grew up, and that’s a confidentiality question
Quick correction to how this is being reported. Voice isn’t an Opus 5 feature. It shipped on its own, July 23, as a Claude app update rather than anything to do with the model.
It’s also the more consequential release for your firm.
Until then, voice ran everything through Haiku, whichever model you were using in text. Now it runs on Opus and Sonnet, reaches connected tools like Gmail and Slack, and asks permission before touching them. Turn-based, so Claude listens, pauses, answers. Still beta. Best from a phone.
Which makes this the first version that could hold a real conversation about a real matter. That’s the problem.
A partner talking through a matter in the car with the Gmail connector live is generating a written transcript that lands in chat history. Confidentiality surface. Possible discovery surface. And it showed up on every associate’s phone without anybody at your firm deciding it should. Free accounts stay on Haiku with one connector, which is worth knowing if your paralegals are logging in with personal accounts, and some of them are.
Before you switch
Trip Anthropic’s cyber safety classifiers inside Claude.ai, Claude Code, or Cowork and the request falls back to Opus 4.8. You will not always be running the model you picked. Mostly this touches security and investigations work, though it’s exactly the sort of thing that surfaces at the worst possible hour.
Sonnet 5’s introductory pricing dies August 31. September 1 it goes to $3 and $15. Build the budget on the September number.
The tokenizer one is buried in the pricing docs and nobody mentions it. Everything from 4.7 forward turns identical text into roughly 30 percent more billable tokens than the older models did. Which means if you’re holding this year’s spend against last year’s and drawing a conclusion about efficiency, you’re not. You’re measuring a tokenizer change and calling it a trend.
Monday morning
Sonnet 5 as the default, Opus 5 for anything that reaches a client or a court without a second read. That’s the whole policy.
Write the voice rule before somebody needs it. One paragraph covering whether voice gets used on client matters and which connectors stay live when it does. Enterprise owners who want it off entirely have to ask Support, since there’s no toggle for it.
Pick one recurring task and run it both ways for a week. Count the fixes, not the minutes.
I’ll be writing one of these again in six weeks. The question underneath it won’t have moved.


