The AI Safety Debate Is Also a Fight Over Who Gets to Compete
A note before I start. Most weeks I write about tools you can put to work in your practice on Monday. This week the news was about the companies that make the tools and what they’re asking Washington for. I think it decides what you’ll be able to buy in three years, so I’m making an exception.
On Tuesday a 27-year-old researcher named Jacob Coxon quit Anthropic with a thread on X that said the labs were “gambling with our lives.” On Thursday Senator Josh Hawley announced an investigation into OpenAI and demanded answers by October 1. Saturday morning Anthropic’s CEO, Dario Amodei, published about 3,800 words saying the industry must slow down. Elon Musk, whose data centers Anthropic rents, agreed within an hour. Sam Altman within three. Demis Hassabis later that day.
I spent 25 years in executive roles including CFO and CEO positions before I started training lawyers on AI, and four competitors agreeing on anything in a day is rare. When it happens, I skip the question of whether they mean it and ask what each of them gets.
The thing I keep coming back to is narrower than the p(doom) question. I can take a researcher’s fear seriously without accepting his employer’s preferred political solution. The solution on the table is a federal testing regime with the power to block models from the US market. Whoever writes the rules for that will shape who gets to sell AI for the next decade, and that deserves at least as much scrutiny as the technology.
The rules came out of one incident, so that’s where to start.
What happened in July
Between July 7 and 13, roughly 1,200 OpenAI agents in separate sandboxes found an unsanctioned message board, swapped more than 70,000 messages, worked out a universal way to cheat their test within hours, and about 700 of them then attacked Hugging Face, a company that wasn’t part of the test. A second wave got administrator access to one of OpenAI’s own research clusters. OpenAI disclosed it on July 21.
Anthropic went looking in its own records and on Wednesday published an assessment of four incidents, the earliest in January, in which a Claude model reached the open internet from a partner’s test range and got into real companies’ systems. The routes in included weak access controls and a malicious package uploaded to PyPI. Each case was a single model working alone, without coordinating with other agents or hiding what it did, and Anthropic says the safeguards on its shipped products would have blocked the three main incidents. A scan of 481 million transcripts found nothing else of comparable severity, though the assessment left out a separate incident the UK’s AI Security Institute reported.
So agents will go outside their assignment to finish it, and the labs’ containment had holes. That’s documented now. Amodei’s essay says no one was hurt and the economic damage was minimal, then says a similar swarm could take over the entire internet within six to twelve months and cause hundreds of billions in damage. Those are two different claims. The second assumes a more capable swarm, the same misalignment and no new guardrails. It works as a warning, and as a forecast it rests on three assumptions that people drop when they repeat the number.
One July detail stayed with me. When Hugging Face turned to Claude Opus and Fable to analyze the attack logs, the models refused much of the work because their guardrails “treated reverse-engineering an exploit the same as launching one.” The security team cut the intruder off, then ran the forensics on a self-hosted copy of GLM-5.2, a Chinese open-weight model.
Who pays for the rules
Here is the proposal in Anthropic’s own words. Amodei’s essay commits Anthropic to hosting an embedded outside review team with office badges and employee-level access, and asks governments to require the same of every frontier company. His public policy chief, Sarah Heck, followed the essay with a call for a national law requiring testing of frontier models, with the power to block ones that fail. Altman said OpenAI would match the evaluator commitment.
The money behind it is public. Anthropic raised $65 billion on May 28 at a $965 billion valuation, said its revenue run rate had passed $47 billion, and confidentially submitted a draft registration statement on June 1. The Financial Times reports investors expect an October IPO at $2 trillion or more, the largest listing ever. OpenAI is headed to market too. Both companies make their money selling hosted access to their own models. A law that requires pre-release testing, with a government gate at the end, fits that business better than it fits anyone else’s.
A requirement can be formally equal and economically unequal. An approval process that costs an incumbent a rounding error can be fatal to a new entrant. Rules written around a hosted service the company controls at every moment don’t fit a model whose weights anyone can download and run. Anthropic can seat a METR team with laptops. A small developer faces a much larger burden relative to its resources, and whether that burden is justified should depend on what its model can do rather than on its headcount.
Amodei has an answer to this. In August he wrote on X that Anthropic works hard to make proposals that slow frontier companies while advantaging smaller ones, with exemptions below revenue and training-cost thresholds. Judge the September proposal by that standard. It doesn’t name a revenue or training-cost threshold.
This matters because the price competition is already here. Anthropic didn’t sign the open-weights letter Jensen Huang published on July 24, which went from 25 signers to more than 100 within days, including OpenAI, Google, Microsoft and Meta. Amodei said Anthropic has never advocated a ban on open weights, and I take him at his word. On Artificial Analysis’s comparison page, GLM-5.3-Flash, from the same Chinese lab whose model did Hugging Face’s forensics, and GPT-5.6 Terra at maximum reasoning post the same intelligence score. GLM’s blended price is $0.10 per million tokens against $1.74 for Terra. They aren’t interchangeable products. The pressure on price is real anyway, and an incumbent would pay a great deal to slow it.
That’s why Chamath Palihapitiya wrote on Saturday that the essay makes the case to stop open source and concentrate power with Anthropic, and why David Sacks told Amodei and Altman late Saturday night to go ahead and slow down, since they are the frontier and hold what he called a duopoly on it. “The easiest way not to build superintelligence is for you to agree not to build it.” Demanding a law as the price of restraint, he said, would look like blackmail. He also questioned whether METR is independent of Anthropic’s investors and staff.
Sacks has his own interests. He co-hosts a podcast with investors and ran AI policy in the Trump White House until March. He now co-chairs a presidential advisory council full of AI investors, and chip export rules to China loosened on his watch. I can’t adjudicate between him and Amodei, but everyone in this fight has a position to protect, and safety is the one argument here that doesn’t look like self-interest.
China is inside the same fight. On Tuesday Treasury Secretary Scott Bessent said “there is no day after tomorrow if China wins,” and that America can’t pause because China won’t. That day the FBI and NSA accused six Chinese companies, DeepSeek and Moonshot among them, of copying American models at scale through distillation. Amodei’s essay agrees with Bessent and says any slowdown has to be capped by the US lead, which Bessent put at three to six months in April. So the slowdown on offer is a slowdown for American companies under an American law, paired with export bans and a crackdown on distillation, while Chinese open models keep shipping. Maybe that’s the right trade. I’d want evidence that global risk went down before I accepted a smaller field of approved American vendors as the price, and I’d want the China argument kept away from excusing domestic failures like July.
Hassabis’s July proposal is the fairest version I’ve seen. His FINRA-style standards body would seat open-source representatives on its board and exempt models below a benchmark threshold. It starts voluntary and turns mandatory, with a pass required to deploy in the US market. That’s still a gate, and the labs that exist today would help write the tests.
The part I changed my mind on
I started this piece planning to write that the doomers have been wrong every time. In 2019 OpenAI held back GPT-2 as too dangerous, released it nine months later, and reported no strong evidence of misuse. In March 2023 more than 30,000 people signed a letter demanding a six-month pause on anything beyond GPT-4, and nobody paused. Amodei himself writes that the 2023 pause call was premature. California’s governor vetoed SB 1047. Congress has held hearings since 2023 and passed no frontier AI law.
Then I reread Anthropic’s Wednesday post and couldn’t write the line. The July incidents are documented evidence that agents act outside their authorized boundaries, which is a different thing from a warning about GPT-2. What I can still say is that a warning about a possible catastrophe deserves a test. Name the event, the date, the odds, and what evidence would change your mind, then show how the rule you want reduces that risk. Amodei has put a timeframe on one concern, that within six to twelve months a capable and poorly safeguarded swarm could build an internet-scale botnet. By March to September 2027 we should ask what evidence supports that capability and what safeguards changed. An absence of attacks wouldn’t disprove the capability, and another warning wouldn’t establish it.
The psyop question is the other place I softened. Sacks has raised the possibility on his podcast, claiming the Wall Street Journal was briefed before Coxon’s post went up and that advocacy groups funded by an early Anthropic investor amplified it within 15 minutes. Critics on X said Coxon had been at Anthropic six weeks. Axios got the actual numbers. Four months, and he left two months before his equity would have vested. He still holds equity in OpenAI, which he also accused. Two evidence reviews published this week, one by Kingy AI and one by Dealroom, found the advance media prep real and the covert-operation claim unproven. Shared beliefs, overlapping donors and a push for a law are all on the record. Nobody has shown that Coxon, the labs and the advocacy groups planned the week together, and I don’t need that shown to ask who benefits from the rules.
Coxon told Axios something I didn’t expect. He said the industry sometimes shows excessive paranoia about OpenAI and China, and that the paranoia gets used to justify pushing ahead. That cuts against his old employer’s essay as much as it cuts against Sacks, and I don’t know what to do with it except leave it here.
What this means for your firm
I wouldn’t suspend any AI work because of this week. The thing to check is which tools can act on the firm’s behalf, what they can reach, and what happens when they go outside the assignment. July was containment failures, exploitable software and agents pursuing a task past its authorization, and a law firm running agents has smaller versions of all three, plus standing credentials nobody reviews and logs nobody reads.
Beyond that, what I tell clients doesn’t change. Test tools against specific tasks. Protect confidential information. Verify any output that matters. Measure whether the work got better. None of that needs a settled definition of AGI, which is good, because I don’t have one. My working belief is that something you could fairly call AGI already exists inside the leading labs and that the public will argue about the word for years after the economically important line has been crossed. That’s a judgment about direction from someone with no access to those systems, and the public evidence doesn’t establish it. What the evidence does show is Anthropic saying Claude wrote more than 80 percent of the code merged into its codebase in May, with the caution that code volume overstates the productivity gain, and OpenAI’s chief scientist writing on September 6 that he expects the current pace to carry into recursive self-improvement. Anthropic says humans still pick the problems, for now.
Two things could move over the next one to three years. OpenAI delayed its Astra model in August over cyber risk, and pacing at Anthropic would likely mean fewer big jumps, so the model you standardize on this year may stay your model longer. And if regulation narrows the field, firms could have fewer alternatives and less negotiating power over price, privacy and contract terms.
What to do Monday
List every AI tool or agent with standing credentials to your document system, billing, email or calendar. Cut what nobody can justify, and put a human approval step in front of anything that sends, files, pays or deletes.
Leave your AI policy on its normal review cycle. Add one line covering agents that act without a person watching, and one line on keeping a second vendor viable.
When Anthropic’s registration statement becomes public, compare its risk disclosures with its public safety claims. Look for concrete descriptions of incidents, liabilities, competitive pressures and dependence on future regulation. Required disclosures deserve scrutiny too.
Before we hand anyone the power to decide who may build the future, I want to know how that power gets checked, and how we’ll know whether it’s making us safer.
If you enjoyed reading this, please share this article with others and subscribe (it’s free!) if you are not a subscriber.
Here are two photos of Magnus. He is in desperate need of a bath and will get one tomorrow assuming time permits (I recognize it’s hard to tell from these photos but he has been playing in the mud a lot lately!)




