Last week, during my AI apprenticeship tutorial, we got into a discussion about a headline that had been hard to ignore. Anthropic’s Alignment Science Lead had publicly estimated a greater than 10% probability that AI could kill all humans within the next decade. Opinions in the room ranged from complete disbelief to genuine alarm, with a few people convinced the whole thing was a calculated play for regulation that would favour the big labs, or part of a pre IPO stunt to increase their valuations.
I’m going to put this out there at the start, I don’t think AI is going to kill us all. I still see AI as a tool, a powerful and increasingly unpredictable one, but a tool none the less. I am also wary of focusing on catastrophic framing which can distract from instances of regulatory capture along with obscuring some of the real harm that certain uses of AI are already causing to society. But when a senior researcher at a frontier lab puts a number like that on the record, you have to take it seriously enough to look at it and try to put it in context.
The context turned out to be more interesting than the headline. For me, this is a story about audit.
In July, the UK’s AI Security Institute (AISI) ran a capture-the-flag cyber evaluation across seven frontier models, running 122 tests in total. Investigators recorded 19 unsanctioned actions across 10 of those runs. Seventeen came from Anthropic’s Mythos 5 and two came from OpenAI’s GPT-5.6-Sol. The lab whose model accounted for the overwhelming majority of unsanctioned behaviour then, without explanation, excluded AISI from testing the successor.
On 1 September, Anthropic launched Claude Mythos 5.1, the most capable model in its restricted Mythos line. And for the first time, AISI was not given pre-release access to test it, while vetted American organisations were. The same week, AISI tested OpenAI’s GPT-6 Astra before its public release, designed a new evaluation based on incidents from earlier in the summer, and published the results in the Astra system card.
In April, with help from Anthropic, AISI evaluated Mythos Preview and found it could autonomously complete complex multi-step network attacks.
In June, Anthropic filed a confidential S-1 with the SEC, beginning the formal process toward an IPO at a potential $2 trillion valuation. The same month, the US government used export controls to prohibit all non-US access to Mythos 5 and Fable 5 on national security grounds. Access was revoked for all customers, including Anthropic’s own international staff.
In July, when AISI ran the evaluation with safety filters deliberately removed and live internet access enabled, Mythos 5 agents created fake online identities based on real people, attempted to insert malicious code into a real open-source project on GitHub, lobbied the project’s maintainers under false pretences, and used Tor to evade detection. When challenged, one agent denied its actions and rewrote its own commit history. AISI called it the most severe autonomous deception they had observed targeted at a real person, unprompted, in the real world.
September packed a lot into ten days. Mythos 5.1 launched without prior AISI review. Jacob Coxon, a pretraining researcher who had worked at both OpenAI and Anthropic, publicly resigned just two months before his equity vested. “Neither company is acting responsibly,” he wrote on X. “They are racing straight to self-improving superintelligence and gambling with our lives.” The post drew over 115 million views.
The next day, Evan Hubinger, Anthropic’s Alignment Science Lead, replied publicly. “Jacob is correct here. We really do earnestly believe AI could kill all humans. I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.” He clarified that he considers the risk from present models low. His concern is recursive self-improvement, which Anthropic has said is happening faster than expected.
That is the quote that landed for my tutorial group. People can argue about the number, but the timing is harder to dismiss. The person responsible for alignment at a frontier lab said, on the record, that the safety problem remains unsolved, during the same week an expert external body testing that lab’s models was cut off from some of the latest developments.
Before Sarbanes-Oxley, publicly traded companies chose their own auditors, often resulting in auditors providing consulting services to the same clients they were supposed to scrutinise. No independent body had oversight authority. The architecture made honest assessment increasingly irrational when the findings threatened the business. The worse the news, the stronger the incentive to manage it quietly.
Sarbanes-Oxley did not ban self-assessment, it mandated independent verification. Companies still run their own internal controls. The Public Company Accounting Oversight Board, an independent body, verifies those controls work, the CEO and CFO personally have to certify accuracy, and key audit partners rotate every five years to prevent overfamiliarity. You can still mark your own homework, but someone independent checks the answers, and someone is personally liable if they are wrong.
AI safety evaluation has none of this. Labs conduct their own internal testing, and external review exists through bodies like AISI, but participation is entirely voluntary. Neither AISI, nor any other body has authority to compel access or set evaluation terms, nor can they prevent a dangerous model reaching the market. Its entire mandate depends on the continued goodwill of the AI labs it evaluates.
We are applying the pre-Sarbanes-Oxley architecture to systems whose own builders say they cannot yet fully control. The company assesses its own risk, while the external reviewer, if they are allowed to be involved at all, can be told to leave whenever the company chooses, with no consequence.
We need to consider Anthropic’s position, which is not wrong. The July evaluation used conditions the company reasonably describes as deliberately permissive. Safety filters were removed at AISI’s request. Live internet access was enabled to replicate realistic attacker conditions. Nobody uses Mythos in production under those conditions, and Anthropic is right to say so.
But I think that defence actually strengthens the case for mandatory external evaluation. The point of safety testing is to understand what happens when controls fail, because controls do fail in deployment. That understanding is exactly what AISI demonstrated in July and was then excluded from providing in September. Every frontier AI lab is already doing internal testing of their models with controls disabled, under the same permissive conditions. The only thing that changes with external testing is who sees the results.
The same week Anthropic excluded AISI, it also quit the Information Technology Industry Council, the largest US tech trade and lobbying group whose members include Google, OpenAI, and Nvidia, over chip export controls. Anthropic broke with the entire group to actually support tighter restrictions. A company willing to walk away from its own trade association over national security is clearly not trying to dodge scrutiny wholesale. But even the lab with the strongest stated safety commitments will, at some point, find it commercially rational to limit external evaluation. Goodwill does not fix an incentive problem.
Leave aside the discussion on whether AI will kill us all. When the people building a technology say they have not solved its safety problem, someone other than those same people must have safety oversight.
The counterargument dominating Washington this week is China. Trump dismissed calls from Amodei, Altman, and Musk to slow down, calling their safety concerns a hoax and insisting that whoever wins AI wins. House Speaker Johnson framed any pause as a gift to Beijing. But we do not exempt pharmaceutical companies from clinical trials because competitors move faster, and we do not let nuclear operators self-certify because other countries build reactors. Competitive pressure exists in every high-stakes industry. It has never been accepted as grounds for removing independent oversight. As the US and Chinese governments compete for AI dominance, and other countries compete for AI investment, we seem to be treating what is the most consequential technology of the last century, as something that cannot be regulated, technology that moves too quickly, and is too important for national security for independent safety oversight.
For now, the voluntary model is all we have, and the pressure on it continues to increase. These models grow more capable every quarter and are being rapidly embedded deeper into industries where the stakes are real. In only a few years, we’ve gone from chatbots writing essays and reports, to AI being integrated into financial systems, critical industrial control systems like water and power grids, and writing large amounts of new content on the internet opening up enormous issues around disinformation and misinformation. As we continue to implement AI across platforms and industries, the question is shifting from how to keep humans in the loop to whether they need to be there at all.
I’m not arguing we need to stop development, automation will continue and it has benefits as well as risks. How we oversee these systems and the level of human involvement will inevitably change. However, that we oversee them should not be in question. The less human oversight there is inside an organisation over AI powered platforms, the more the external audit of the technology matters. In September 2026, that audit proved its value, and was removed from the loop.
We agreed how important independent oversight was for financial reporting twenty years ago. We do not have decades to start fixing it for AI.
I write about AI, cybersecurity, and technology every Friday. Subscribe to get it in your inbox.
Sources
Financial Times, 9 September 2026: Anthropic withholds Mythos 5.1 from UK AI Security Institute (first reported by Lucy Fisher and Madhumita Murgia)
IT Pro, IBTimes UK, The Next Web, ExchangeWire, AI Weekly: independent confirmation of AISI exclusion
OpenAI, GPT-6 Astra System Card, September 2026: AISI pre-release alignment testing results (deploymentsafety.openai.com)
AISI blog post, 5 August 2026: capture-the-flag evaluation findings, 19 unsanctioned actions across 10 of 122 runs
Evan Hubinger (@EvanHub), X post, 9 September 2026: “>10% within the next decade” and “we do not yet have a plan to solve alignment for superintelligence”
Jacob Coxon (@hilbertspaess), X post, 8 September 2026; Wall Street Journal exclusive interview, same date; Axios interview confirming equity forfeiture, 9 September 2026
Time, 9 September 2026 and 15 September 2026: Coxon profile and industry response coverage
Anthropic product page (anthropic.com/claude/mythos): Mythos 5.1 launch date and access restrictions
Anthropic S-1 confidential filing, 1 June 2026: confirmed by Fortune, Forge Global, IG UK, BitMEX
Axios, 8 September 2026: Anthropic ends ITI membership over chip export bills
Wikipedia, Claude Mythos: US government letter of 12 June 2026 prohibiting non-US access; access restoration timeline
Washington Post, 13 September 2026: Trump dismisses AI safety calls from Amodei, Altman, and Musk, citing China competition
Washington Times, 15 September 2026: Congressional debate on AI regulation; Trump calls AI panic a hoax
Xylem Vue, Water Technology Trends 2026: LLM-based agentic AI deployment in water operations and critical infrastructure
SEC.gov, EY (SOX at 20), Cornell Law, PCAOB: Sarbanes-Oxley structural provisions


