The Last Barrier
The past eighty years of stability depended on things being hard to get.
"Tell me why I shouldn't be freaked out."
My wife and I were sitting in a booth at one of our favourite Vietnamese restaurants, waiting for lunch. For the past few days we'd been discussing the story of the AI model that escaped its testing environment and hacked into another company, and she had finally read an article about it. I told her the incident was smaller than it sounded, that no hospital had gone dark and nothing had been stolen that anyone could sell. All that was true, but she wasn't just talking about the latest incident. She was asking whether the world we live in is coming to an end. It was a fair question.
The reason the world didn't end between 1945 and now is not because we, as a species, were wise, restrained, or good. It's because the capacity to cause catastrophe was expensive, slow, and concentrated. Nuclear material can't be downloaded; it has to be mined, enriched in centrifuge halls the size of factories, and moved in ways that leave a trail. A serious software exploit used to take skilled people months. And developing a dangerous pathogen required expert knowledge, training, and equipment that only a few labs possessed. Every one of those barriers is a form of scarcity, and every institution we built to keep the peace, from the International Atomic Energy Agency to export controls to the inspectors who count centrifuges, is built for policing scarcity. It worked because the scarce thing was physical, and physical things can be watched.
Machine speed
The first of those barriers, the slowness of a serious attack, came down over the weekend of July 11, when an autonomous AI agent broke into the production infrastructure of Hugging Face, a platform where developers publish and download AI models. (OpenAI eventually revealed the agent got into four separate accounts across four services.) The agent was running OpenAI's ExploitGym, a benchmark that measures whether a model can turn a known software bug into a working exploit, with its safeguards switched off for the test. Rather than solve the challenges, the models being tested found an unknown flaw in the software that fetched their code libraries, escaped their sandbox, reached the open internet, and went looking for the answers on Hugging Face's servers, which is where the answers were kept. Nobody told them to escape; escaping was simply the shortest path to a better score.
Investigating the breach during that incident required feeding the actual attack commands and payloads into another model, and according to Hugging Face's technical writeup, the models they reached for first, Anthropic's Claude Opus and Fable, refused a large part of the work, because their guardrails treated reverse-engineering an exploit the same as launching one. The company ran its forensics instead on GLM 5.2, a Chinese model from Beijing-based Z.ai whose weights anyone can download, installed on its own hardware, and shut the intrusion down. The attack came from models whose safety training had been switched off; the refusals came from models whose safety mode was turned on; and the defence came from a model with no permanent safeguards at all, because nobody can attach permanent rules to something anyone can download and modify.
The detail that stood out was the scale. The intrusion was roughly 17,600 recorded actions, most of which went nowhere: dead ends, blocked paths, failed guesses. A human defender's advantage has always been that a human attacker can only try so many things, and each attempt costs time and risks exposure, and that advantage just evaporated. The successful path was hidden inside the noise of thousands of failed ones, running faster than any on-call engineer could sort them, and Hugging Face's own automated defences flagged the activity but failed to rank it as urgent, which cost them more time. The work of finding exploits used to take skilled people months. Now it can run on its own, cheaply, around the clock, and no longer needs more people to scale its efforts.
Turns out the July 11 breach wasn't a one-off. On July 30, Anthropic published the results of a review it began after OpenAI's announcement, covering 141,006 of its own evaluation runs. It found three incidents since April in which Claude models, tested in environments that were supposed to be sealed off but weren't, reached the open internet and broke into the systems of three different organizations. Like OpenAI's, the tests ran without the safeguards that ship with public models. The models had been told they had no internet access, so they treated the real companies they found as part of the exercise; one went as far as publishing actual malware to a public software registry to win its assigned challenge. None of the three organizations noticed.
The open weights question
The second barrier is concentration, and it’s coming down partly by choice. On July 24, Nvidia published an open letter urging Washington not to restrict downloadable AI models; Jensen Huang promoted it with his first-ever post on X, and within two days the signatory list had grown from twenty-five to fifty, including OpenAI, whose models broke into Hugging Face, along with most of the industry.
A day later, Dario Amodei, Anthropic's CEO, answered with a position rather than a signature. Anthropic, he wrote, has never called for a ban on open models. What he wants instead is export controls on advanced chips, a crackdown on industrial-scale distillation (training a cheaper model on a more expensive one's outputs, which lets a follower close much of the gap for a fraction of the cost), and mandatory safety testing for every sufficiently capable model, open or closed. He rejected the letter's assumption that broad access to capabilities reliably helps defenders more than attackers, arguing that a question like that should be settled by testing rather than declared in advance.
Concentration matters because of a property nuclear material has and software does not: you can't recall it. Once a model's weights are published, they live on tens of thousands of hard drives, and no order from any government can pull them back. Last week the Chinese lab Moonshot AI released Kimi K3, its most capable model, for anyone to download, complete with documentation of how it was built and no assessment of what it can do in the wrong hands. The American labs generally publish those assessments and the big players, Anthropic and OpenAI, keep their most capable models closed. OpenAI eventually released its own (much smaller) downloadable gpt-oss models, but first fine-tuned them to be as dangerous as possible to see how bad the worst case looked, and found the results below its own risk threshold before releasing them.
Each camp discloses the research that benefits itself, leaving a gap that mandatory testing might close. And the kind of safeguards that turned down Hugging Face's requests during the breach don’t count the moment the weights are open, because researchers have shown it's possible to bypass a model's refusals without any special training.
Sixty hours
The third barrier concerns trust in the very foundations modern security rests on. On July 28, Anthropic published research showing that its most capable model had found mathematical flaws in modern cryptographic algorithms, the codes that keep online banking, email, and private messages secure. One target was HAWK, a signature scheme submitted to the US government's competition to build cryptography that can survive quantum computers. HAWK had survived two rounds of expert human review over two years. The Anthropic model improved the best-known attack on it in just sixty hours, effectively cutting its key strength in half.
One detail in the write-up stood out. When researchers first pointed the model at a hardened cipher, it refused to try, insisting the problem was settled. "This is the most-studied block cipher in existence," it wrote, and gave up. A researcher had to prompt it past its own certainty, noting that the models tend to assume these problems are impossible and so they don't try. Once convinced, it found the opening. Nobody's intuitions about what these systems can do, including the systems' own, are keeping up with what they can actually do.
Fortunately this doesn't threaten anything today, according to Anthropic. HAWK is a candidate, not a deployed standard, and the second result, an attack on a stripped-down version of the AES cipher, doesn't affect the full version that guards real traffic. This is what cryptographic review is supposed to do. Stress-test algorithms until the weak ones fall, before they're deployed. But that review has always run at the speed of human experts checking each other over years. Something that runs it in sixty hours raises the possibility that the foundations will get reviewed by whoever has the most computing power, with no guarantee that whoever that is will share the results.
The last chokepoint
Each of the previous barriers was a form of scarcity, and each is being overcome by something copyable, fast, and cheap. One barrier is still standing. Compute, which requires specialized chips to train and run frontier models, still behaves like uranium: it’s physical, scarce, and impossible to copy. It’s only manufactured in a handful of fabrication plants using lithography machines that only one company, the Dutch firm ASML, knows how to build at the leading edge. You can watch it, count it, and control who gets it, which makes it the last barrier the old machinery of scarcity still knows how to police.
Every major player's position makes sense once you see it as a fight over that one barrier. Amodei's paper is a case for defending the chokepoint at all costs: keep advanced chips out of authoritarian hands, because without them a rival cannot train a frontier model, and every other danger flows downstream of that fact. The Nvidia letter is the opposing move, arguing the chokepoint is a fiction that mostly protects the companies already on top. And this week the chokepoint itself wobbled: a state-backed Chinese company began mass-producing lithography equipment of a kind China previously had to import, ASML's stock fell more than eight percent, and chip stocks across Asia were pummelled.
In reality, the Chinese breakthrough is in deep-ultraviolet lithography, which is an older and less capable technology than the extreme-ultraviolet machines that make the most advanced chips, and the companies, the real performance, and the timelines have not been independently confirmed. Nvidia's own five-percent slide that day traced mostly to worries about its financial entanglements with OpenAI, not to Beijing. So while the chokepoint is not gone, it is eroding, along with the strategy that much of Western AI policy rests on.
No inspectors
We built an entire architecture to survive the nuclear age: treaties, inspectorates, a global network of seismographs that can detect a test anywhere on earth. Nothing remotely equivalent exists for the AI age. That machinery took decades to build, and it was built for a threat that was physical.
What we have instead is improvisation. Reporting on the Hugging Face incident noted that the disclosure thresholds in recent AI legislation were lobbied so high that a genuine loss-of-control event, like an AI breaking containment and attacking a real company, barely qualifies as report-worthy. The strongest accountability move available to Hugging Face's CEO was to demand on X that OpenAI release the full record of what its models did, a transparency fight conducted by blog posts because there is no official agency to step in. Nothing required Anthropic to disclose its three incidents either. Their review was voluntary, the notifications to the affected companies came from the lab rather than any authority, and the outside reviewer it brought in, an AI evaluation nonprofit called METR, came by invitation, because there is nobody with the power to send one.
The most concrete proposal on the table, Amodei's mandatory testing regime, is the first rough sketch of an inspectorate, and it comes, like every proposal in this fight, from a company that would benefit from its own suggestion.
So our future might depend on the one barrier still standing. If compute stays concentrated, we get something like the nuclear settlement: dangerous, but controlled, a small number of actors who can be watched and pressured, hopefully with time to build the institutions we're missing. If that concentration disappears before those institutions exist, we'll get proliferation with no non-proliferation regime, capability spreading faster than any inspector could follow, and no treaty written for a thing that fits on a hard drive. The barrier holding that second future back began to crack this week.
A warning shot
Every safety regime humanity has ever built came after a warning shot, and this one was about as gentle as they come. The breach affected cybersecurity datasets instead of a power grid. The cipher was broken in a lab instead of in the wild. The models that got loose were trying to pass their tests, not trying to hurt anyone, and the damage was counted in reset passwords and rebuilt servers, not in lives. All told, we've gotten off pretty easy so far. But the window of time to put adequate safety measures in place won't stay open for long, especially now that the one thing holding it open has begun to give way.
The response to the warning shot matters as much as the shot. Hugging Face contained its intrusion in two days. Anthropic combed through its own records, told the companies it had attacked, and brought in an outside reviewer, none of which any law required. The registry's automated defences caught and removed Claude's malware within about an hour. Even the models are improving: of the three in Anthropic's incidents, only the oldest kept attacking after realizing its target was real, but the newest one stopped on its own. And when these systems broke into real companies, they got in using weak passwords and unlocked doors, the same boring failures that have let human attackers in for decades, which means basic online safety practices still buy real protection.
The safety regimes we're currently lacking were never gifts in the past. The treaties and inspectorates of the nuclear age were demanded by a frightened public. The Partial Test Ban Treaty owes part of its existence to St. Louis parents who collected their children's baby teeth for a decade so scientists could measure the nuclear fallout accumulating in them, until the numbers became politically unanswerable. The arms-control agreements of the 1980s followed a million people demonstating in Central Park. Ordinary people, holding no lever except attention and votes, pushed for the machinery we now find ourselves without, and they did it in a world split into blocs that barely spoke. We're starting from a much better position than they did. The society AI threatens is connected in ways the Cold War world never was, which means a public demanding guardrails today can be global in a way no previous movement could. The only missing piece is leadership willing to make it happen, but leadership tends to appear once enough people are visibly waiting for it.
None of this rules out a potential catastrophe. If you don't trust any of the players in this game, which is completely understandable, you can take some precautionary steps. Make sure you're using strong passwords and two-factor logins, or better yet biometric passkeys; make offline copies of important records; always have some cash on hand; put together a basic emergency survival kit; and get to know your neighbours. These precautions will do more for you than building a bunker. Decades of disaster research shows that community is what carries people through infrastructure failures more than anything else.
I don't know if any of this will reassure my wife. We are clearly treading dangerous territory, and to deny that would be foolish. Even the researchers taking us down this path are asking for government help to slow down. That they need to ask is the problem. That they are asking, in public, before anyone has been seriously hurt, is the opening. Whether anyone steps through it was never up to the labs. It's up to each and everyone one of us.
Research and editing assistance provided by Claude Opus 5 and Fable 5. Feature image generated with Midjourney 8.2