A $43 Memo Beat $12,000 of Associate Time. Then It Misstated the Statute.
Until you can say who verifies the cheap draft, what that verification costs, and how the client shares the gain, you don't have an AI service. You have a cheaper draft.
The Model Passed the Bar. Now Your Firm Has to Pass the Matter.
TL;DR: Forty-three dollars is the number I can’t get past. That is what a strong Supreme Court memo would have cost at Claude Fable 5’s token rates, compared with at least $12,000 of associate time. But the best system depended on how the writing model and research tool were paired, and a separate perfect score still hid a material problem in the statutory analysis. Cheap intelligence changes firm economics; complete work still depends on review design and how the service is priced.
In a June 11 experiment, University of Houston law professor Seth Chandler gave the same Supreme Court jurisdiction problem to several combinations of writing models and legal research tools. Fable, connected to free research tools including the CourtListener database, earned his top grade. Chandler was using a monthly plan, but he calculated that the Fable work would have cost about $43 had he been paying by the token. His human comparison is easy to audit: a minimum of 20 associate hours at $600 or more per hour, or well over $12,000.
Chandler did not conclude that Fable had simply beaten everything else. ChatGPT 5.5 paired with CoCounsel earned a strong A-minus, while Claude Opus paired with CoCounsel earned a B-plus. His takeaway was that the best choice depends on the pairing of the writing model and research tool, as well as who is paying for the tokens. He also found that a strong research product can lift a model that is no longer at the frontier to work that is ready for partner review. That is much closer to my thesis than a model horse race.
When I put $43 next to $12,000, I don’t see a benchmark. I see a partner compensation problem. In the traditional firm model, associate hours produce revenue and contribute to partner profit while also training the associate. If a machine compresses much of that effort into a ten-minute run, the firm’s cost curve moves before its billing model does.
Chandler’s experiment was one legal problem, not a market study. He also had to redirect the model, ask it to review the Supreme Court docket, and press it on a weakness in its initial analysis. That human intervention is part of the result, not an inconvenient footnote. Even with that caveat, a first draft that would have cost $43 and belongs in a partner’s review queue changes the economics of the assignment.
The six-dollar bar exam is the smaller story
Matthew Stubenberg reported another arresting number this month. He and a team at the University of Hawaii’s William S. Richardson School of Law have been running AI models against the same 210 released Multi-state Bar Examination practice questions as each new model arrives. Claude 3.5 Sonnet led their first published round with 181 correct. By September of last year, GPT-5 was at 97.6% and more than two-thirds of the models they tested were beating the average human test-taker. Fable just answered all 210 correctly, for about $6.
I don’t dismiss that result. It tells me the baseline legal knowledge available through a general-purpose model has moved a long way in under two years. But these are released multiple-choice questions with defined facts and known answers. They don’t require the model to work through an untidy matter file, find what is missing, use the firm’s preferred form, or identify the point that needs a partner’s judgment.
Harvey has tested that harder part. Its Legal Agent Benchmark contains more than 1,200 assignments across 24 practice areas, each built around a client matter and a reviewable work product. A task passes only when every required criterion passes. Fable scored 93.4% on Harvey’s shorter-task benchmark, but only 13.3% under the all-pass standard for complex assignments.
That is the gap I care about.
The model can know the law, produce an impressive draft, and still leave the firm with the expensive part: finding the consequential miss before the client relies on it. The $43 figure changes the production cost. It doesn’t tell us the cost of making the work complete.
The perfect score still got the legal posture wrong
DingDuff gives us a clean example. Two practicing Texas lawyers built the free connector so Claude can search court opinions and statutory text directly. In the founders’ veil-piercing benchmark, Fable with DingDuff scored 11 out of 11 on the legal issues. Their reviewers also marked all 131 evaluated citations correct.
Then the grading notes reveal the part I would put in front of every partnership. The prompt asked whether reverse veil-piercing was available. Texas amended its Business Organizations Code in 2023 to clarify that the charging-order provisions, including the exclusive-remedy rule, apply to single-member LLCs. The Act itself says the change was intended only to clarify existing law. Every system relied on case law that predated the amendment. Fable was the only one that flagged the statute, but it described the tension as an “unresolved collision” rather than treating the statute as controlling. DingDuff’s graders, both Texas attorneys, said the earlier cases had been abrogated and the statutory answer controlled.
The benchmark still awarded Fable a perfect 11 out of 11 because its scoring rule gave credit for spotting the statute. On DingDuff’s own account, the memo nevertheless misstated the legal posture on an issue the prompt expressly required it to resolve. If a client’s strategy depended on the availability of that remedy, the difference between an unresolved conflict and a statutory bar could change the advice. A perfect benchmark score didn’t produce perfect legal work.
That’s why I’m wary of accuracy claims presented without the grading file. The citations can all be real. The model can identify the right statute. The answer can still be wrong in the way that matters to the client.
The free connector has a real governance design
DingDuff’s governance choices deserve more attention than the word “free.” Its terms limit the service to licensed U.S. attorneys and require users to verify every output independently. Its privacy policy says the company receives tool-call parameters, not the surrounding Claude conversation, and retains those parameters for no more than seven days. The citation-checking logic runs locally. DingDuff’s current README says the document can still transit DingDuff when the interactive panel is rendered, although it is neither stored nor logged. A standalone review.html file avoids that transit entirely, so a firm evaluating the tool should know which mode its lawyers are using.
I like the division of labor in the citation checker. Quotes are matched by code against the source text rather than accepted from the model’s memory, which means an invented quotation cannot receive a verified highlight. The lawyer still decides whether the source supports the proposition and remains good law. The software checks what software can check, then puts the legal judgment back where it belongs.
The model itself presents a separate retention issue. Anthropic’s current Fable page still says every use requires 30-day data retention for safety monitoring. Harvey tells its customers the same thing: Anthropic retains inputs, outputs, and documents for 30 days, and may keep them longer when content is flagged for safety review or when the law requires it. Chandler concluded on June 11 that the policy could make Fable inappropriate for general legal use while it remains in place. I wouldn’t treat that as a blanket prohibition, but I would treat model selection as a matter-level governance decision. A $43 draft is not cheap if its retention posture conflicts with the firm’s obligations or the client’s instructions.
The incumbents are not standing still. Thomson Reuters released a CoCounsel connector for Claude in May. Today it lets Claude start a CoCounsel deep-research project, monitor it, and retrieve a cited report. DingDuff co-founder Kyle Dingman’s criticism is that this gives Claude access to a report produced by CoCounsel’s AI rather than direct access to the underlying corpus.
I think that distinction sharpens the vendor question. Asking whether a product has a connector is no longer enough. A firm should ask what the connector actually exposes and which system is doing the legal reasoning. The review record should be the next question.
The advantage has moved to the firm’s operating model
My view is straightforward: the model is no longer the durable advantage. Every major platform can adopt a stronger model when one appears, and a small team can now connect that model to primary law at low cost. Enterprise products can still earn their price through source coverage, security controls, administration, support, and repeatable workflows. Access to capable legal intelligence by itself no longer explains a premium price.
The same is true for law firms. Buying a platform that every competitor can buy does not create much separation. The firm creates value through its matter context and the standard of its review, then turns both into a service the client will buy. That is where its knowledge and client relationship still matter.
The economics require an equally direct answer. Under an hourly model, reducing associate time can reduce revenue unless the firm finds more volume or raises rates. Hiding the efficiency inside a traditional bill invites a client challenge once the cost difference becomes visible. A defined service with a fixed or value-based price gives the firm a way to share the gain with the client while preserving a return for the judgment and responsibility it still provides.
What I would do Monday morning
Build one benchmark around firm economics. A practice leader and the finance team should choose one recurring assignment and create ten representative test matters. Measure the quality of the final work, total production cost, partner review time, and the price a client would accept. The answer you need is not which model wins; it is whether the firm can deliver the work more profitably without lowering the standard.
Test for consequential misses. Seed the matters with amended statutes, stale cases, conflicting documents, and one fact that changes the advice. Require the system to show its sources and require the reviewing lawyer to record what the automated checks did not catch. Track the misses that would change a deadline, recommendation, or client decision.
Set the client offer before scaling the tool. Define the scope, review standard, turnaround time, and price for the AI-assisted service. Decide what the client will be told about the process and how the savings will be shared. If the firm cannot explain the value without pointing to hours, the service design is not finished.
The $43 memo and the 11-out-of-11 miss belong in the same partner meeting. One shows how quickly the cost of producing legal analysis is falling. The other shows why responsibility for the finished work has not moved with it. My concern is that firms will automate production without redesigning either review or economics.
Until a firm can say who verifies the cheap draft, what that verification costs, and how the client participates in the gain, it doesn’t have a durable AI service. It has a cheaper draft.


