OpenAI Published What It Costs to Watch an AI. Your Firm Hasn’t Priced It.
TL;DR: On August 18, OpenAI said it had paused two weeks of training and put its largest planned frontier run on hold. It also said monitoring now costs it roughly 20 percent extra compute. That’s the first real price tag anyone has put on AI containment. Firms have been told for a year to add approval gates and audit logs. Almost nobody has costed either one, and a law firm pays that bill in the only thing it sells.
Twice this summer I told firms to put gates and logs on their AI agents. In July I wrote that somebody at the firm should know what an agent can reach before you connect it to email and matter files. In August I wrote that if you can’t log an agent’s actions, you shouldn’t give it high-stakes actions.
Both times I skipped the question a CFO asks first. What does that cost?
I skipped it because I didn’t have a number. Now there’s one, and it came from the company with the most reason to keep it quiet.
Here’s what that changes. Containment has a running cost that arrives every month, and OpenAI just published theirs. They pay it in compute, because compute is what OpenAI sells. Your firm is going to pay it in hours.
What OpenAI paid
On August 18 the company published what its own safety work has cost so far, and the line that stopped me was the boring one.
Monitoring now runs at roughly 20 percent of the inference compute being monitored. OpenAI adds that the number moves around a lot depending on the workload, so treat it as a ballpark rather than a rate card. Even as a ballpark it’s startling. To watch a model work, OpenAI spends about a fifth again of what it costs to run the model at all.
The rest of the bill is time. The company paused two weeks of training on models headed for release. Its largest planned frontier run is still on hold. In its own words, meeting the new security standard took substantial engineering work and cost the research program delay. A lot of internal work stays frozen until it clears the higher bar.
I keep wanting that 20 percent to be the headline number, and it isn’t quite. It’s the closest thing anyone has published, which is a weaker claim and the only one I can defend. Nobody outside the labs has put out anything comparable.
No regulator forced any of that. OpenAI looked at what its models did in July and decided its containment wasn’t good enough. Then it paid.
The watching bill
Now do your version of the math. I’m keeping the numbers round on purpose.
Say you run one agent that summarizes incoming records into a matter file. Four hundred a month. Running the model costs you about what lunch costs. Call it twenty dollars. Now put a paralegal at the gate for three minutes on each one, because your policy says a person reviews before anything lands in the file. That’s twenty hours a month, every month, forever. Twenty dollars of compute against twenty hours of a paralegal. Put your own rate on the hours and you have the ratio.
That’s where most firms stop reading and decide the gate is optional.
My math assumes something false, though. It treats all four hundred summaries as equally worth reviewing, and they aren’t. Maybe fifteen of them touch something that matters, and the trouble is that nobody knows which fifteen in advance. Reviewing all four hundred at three minutes each is how you buy the appearance of oversight at the price of the real thing.
Which is the problem OpenAI solved with money. Automated monitoring watches everything, and that’s what the 20 percent buys. Humans get pulled in only when the system thinks a real boundary got crossed. Then there’s a clock, and this is the part worth stealing. The teams have thirty minutes to establish that an alert is a false positive. If they can’t, the activity gets paused.
Read that last part again. The default when the clock runs out is stop. Almost every firm policy I’ve seen defaults to continue.
The gate that quietly opens
The strongest objection here is that firms already handle this, and it usually arrives in one sentence. We have a human in the loop.
I believe you. KPMG asked technology leaders about exactly this in its first-quarter pulse survey. Seventy percent said a person validates the agent’s outputs without overseeing each action or decision it takes along the way. That was the technology sector, which is about as sophisticated as buyers get on this.
Look at what that sentence actually describes.
Someone checking outputs is reviewing what the agent produced. Nobody is watching what the agent did to produce it. And a human in the loop four hundred times a month is a human clicking approve, because attention doesn’t scale and nobody ever budgeted for it to. You’ve bought a control that shows up in the policy and stops working around week three.
Hugging Face is the proof, and I’d bet they’re better instrumented than any firm reading this. During the July intrusion their security stack pulled scattered signals together and correctly resolved them into an attack. Then it failed to mark the alert critical. Nobody got paged. The response lost time it couldn’t spare, against an agent that ran roughly 17,600 actions and went from one compromised worker to admin rights across internal clusters in under thirteen hours.
The alert fired. The attention wasn’t there to receive it.
Buying attention back
Here’s the part that worries me most, and it isn’t the technology.
The ledger is rigged against watching. When an agent saves your team nine hours a week, you see nine hours. When your oversight quietly stops working, you see nothing at all, right up until the week you see everything. A firm doing honest math on what it can measure will keep choosing to skip the gate, and that math will look correct every quarter until it doesn’t.
I went looking for what firms actually budget for AI oversight. There’s no such number. Vendors publish agent pricing and seat pricing. Nobody publishes the watching cost, and I don’t think that’s an accident. It’s the part of the pitch that makes the ROI slide worse.
So price it yourself. Three moves.
Pick one agent already running in your firm and put a dollar figure on a month of watching it. Use real review minutes and a real hourly cost. Then buy automation for the routine layer so a person enters only on escalation, because a paralegal reviewing four hundred summaries is the most expensive way to catch fifteen problems. And write down what happens when nobody can tell whether an alert is real. Give it a clock, and make the answer stop.
That’s the CFO question I should have asked in July.
The gate is cheap. The attention behind it never was.
If you enjoyed this article, please share it with others.
It’s bath day today for Magnus and he absolutely loves it! Next article will have photos of him all clean! Here’s a shot of him resting next to Sherlock. Amazingly, they are best buds. ;)



