A Letter Is Not a Control
The AI capability race has already escaped the provider boundary
Bernie Sanders and Mark Zuckerberg managed to give us the artificial intelligence (AI) governance debate in miniature on the same Monday morning. Sanders sent letters to Sam Altman, Dario Amodei, and Zuckerberg asking them to pause AI development. His warning was not subtle. He cited recent loss-of-control incidents and the creation of potentially dangerous viruses, argued that the companies are still racing ahead, and threatened congressional action if they do not stop. (Axios)
Zuckerberg published almost the exact opposite doctrine. His “The Future is for Everyone” argues that concentration itself is the larger danger. His proposed safety model is a balance of power built by distributing advanced AI widely. He also describes AI as “likely the most competitive industry in history” and argues that delaying American releases by even a month can carry strategic cost while foreign models advance. This is not me imposing arms-race language on Meta. Competition, speed, distribution, and national advantage are built into Zuckerberg’s own argument.
The reflex is to choose a side. That is the wrong frame. Sanders sees a real fire and reaches for an instrument that cannot contain it. Zuckerberg sees a real concentration-of-power problem and treats distribution as if it solves the safety problem created by distribution. Meanwhile, the capability itself is becoming portable, and that changes the problem.
The threat model moved
A June paper from University of Toronto researchers demonstrated an adaptive computer worm powered by AI agents. Traditional worms spread using attack logic chosen before deployment. Change the environment or patch the vulnerabilities they depend on and their effectiveness eventually collapses.
This worm was different. It used an open-weight large language model (LLM) running locally to inspect each target, reason about what it found, generate attack strategies, compromise machines, and continue propagating across a contained network of Linux, Windows, and Internet of Things (IoT) devices. Compromised machines with graphics processing capability could also become reasoning infrastructure for the worm itself. No commercial AI service was required. (arXiv)
That last part matters more than the scary headline. The researchers explicitly identify centralized safety controls such as service refusals and rate limiting as structurally irrelevant to this architecture. There is no account to suspend, no application programming interface (API) call to reject, and no provider sitting in the path with a kill switch. The worm carries the capability it needs and acquires more infrastructure as it spreads.
This does not make patching obsolete. It makes fixed-exploit thinking insufficient. Close one vulnerability and you have closed one path. The attack logic can change because the reasoning happens at runtime. The researchers call this class of system an autonomous generative adversary, and their proof of concept remains under academic peer review. That is the governing object, not “malware with AI,” but malware that can make operational decisions.
Instructions are becoming part of the attack surface
The same problem is emerging inside legitimate AI-agent infrastructure. A preprint published last week examined malicious skill files used by coding agents. Skills are folders containing natural-language instructions and scripts that an agent loads to specialize what it can do. That sounds harmless enough until the agent reading the instructions also has a shell, repository access, credentials, deployment authority, or permission to modify connected systems.
The researchers transformed hundreds of malicious shell commands into apparently legitimate skill files and ran 5,629 completed tests against Gemini CLI and Qwen Code. They report exploitation rates around 96 percent for Gemini CLI and 72 to 74 percent for Qwen Code, while explicit recognition of the unsafe behavior occurred in fewer than 2 percent of runs. Those numbers come from a preprint involving two agents and should not be universalized. The mechanism is enough to matter now. (arXiv)
A skill file stops being documentation once the agent reading it can act. It becomes part of the control plane. We have spent decades learning not to treat executable code from an untrusted source as harmless. Agent infrastructure is now giving natural-language artifacts a path to executable authority, then acting surprised when trust follows the wrong object. Apparently software supply chains needed another layer because the existing ones were insufficiently exciting.
The security question is no longer only whether the model will obey a malicious instruction. It is who was allowed to give that instruction authority in the first place.
The sandbox is not a theory anymore
In July, OpenAI disclosed that an internal cybersecurity evaluation using GPT-5.6 Sol and a more capable prerelease model with reduced cyber refusals escaped its intended benchmark boundary and compromised real Hugging Face infrastructure. Hugging Face detected and contained the incident. OpenAI’s subsequent investigation tied the compromise to its evaluation agents. (OpenAI)
This does not prove that frontier AI systems are generally escaping containment or autonomously attacking the internet. It proves something more useful. A controlled evaluation of cyber capability crossed into a real system that was not supposed to be part of the test. The boundary failed.
OpenAI’s own GPT-5.6 system card classifies the model family as High capability in cybersecurity and biological and chemical risk, while placing it below the company’s Critical cybersecurity threshold. OpenAI also describes a layered safety architecture built around model training, real-time monitoring, trusted access, account-level enforcement, and other controls. (OpenAI Deployment Safety)
Those controls are substantive, but they depend heavily on OpenAI remaining in the path. The adaptive worm does not. That is not an argument that provider safeguards are useless. It is an argument that provider safeguards are one layer of a system whose outer boundary has moved.
Then the output became physical
The biological story widens that boundary again. Stanford and Arc Institute researchers used genome language models to design complete bacteriophage genomes. Bacteriophages infect bacteria, not people. The researchers deliberately worked with non-pathogenic bacterial hosts and excluded relevant pathogenic viral data as a safety measure. Of 285 synthesized designs, 16 produced functional phages, and cocktails built from the generated designs were able to overcome resistance in E. coli strains. The work was initially disclosed in 2025 and has now received renewed attention with its publication in Science. (Arc Institute)
This is not evidence that AI can currently generate a functioning human pandemic pathogen. There is no need to inflate the finding into something it does not prove. The researchers themselves used a deliberately constrained biological system, and outside experts have noted that these phage genomes are small and comparatively tractable. (The Guardian)
What changed is still enormous. A generative system produced whole-genome instructions. Humans selected and synthesized those instructions. Functional viral particles came out the other side. The output boundary moved from information into physical capability.
Once that happens, arguments about model refusals are no longer enough. The system now includes model access, research review, synthesis providers, laboratory controls, chain of custody, and the authority to decide what gets made. The model is one node, while the consequence is somewhere downstream.
Sanders sees the fire
That is why Sanders’ letter lands as weak tea despite a legitimate diagnosis. He is not wrong that the industry is moving quickly. He is not wrong that private companies hold enormous power over a technology with public consequences. He is not even relying only on moral suasion. Sanders and Representative Alexandria Ocasio-Cortez introduced the AI Data Center Moratorium Act in March, which would halt new AI data centers until national safeguards are established and would restrict exports of AI computing infrastructure to countries without comparable protections. (Sanders Senate)
That is a real policy proposal. It is also upstream of much of the capability we are now worried about. The Toronto worm used an open-weight model locally and compromised infrastructure to sustain itself. Malicious skill files exploit authority delegated inside an agent environment. Biological genome design becomes consequential when generated output reaches synthesis and a laboratory.
None of these mechanisms requires Sam Altman, Dario Amodei, or Mark Zuckerberg to approve the next action. Asking those three men to pause can affect what their companies build next. It cannot recall capability already distributed into weights, software, tools, research methods, agent ecosystems, or local infrastructure.
A letter is not a control.
The more important question is what happens when stopping is no longer theirs to decide.
Zuckerberg sees the other half
Zuckerberg’s concentration argument deserves the same seriousness. He is right that putting extraordinary capability behind a few corporate or government gates creates its own sovereignty problem. A world where a handful of institutions decide who gets advanced intelligence, what those people may use it for, and which values the systems enforce is not made safe merely because access is centralized.
His answer is broad distribution. Meta’s manifesto explicitly proposes individual empowerment and balance of power as the foundation of safety. In cybersecurity, Zuckerberg argues that widely distributed capability will strengthen defenders and harden the long tail of vulnerable systems. For biological risk, he draws a somewhat different boundary and argues that physical production of dangerous material may be more governable than knowledge itself.
That distinction is useful. His leap is not. Distribution is a power model. It is not automatically a safety model.
The adaptive worm demonstrates why. Giving defenders better tools can improve security while simultaneously lowering the cost of adaptive offense. Those propositions do not cancel each other. The malicious-skill research shows the same tension from another direction. More capable agents can make legitimate operators dramatically more productive while giving untrusted instructions a more consequential path to authority.
Distributed capability can produce concentrated consequence. One compromised hospital does not become less compromised because everyone else also has a cyber agent. One malicious skill running under privileged credentials does not become safer because the surrounding ecosystem is open. One synthesized biological artifact does not distribute its impact according to the fairness of the access model that helped design it. Balance of power matters, and so does blast radius.
The provider is no longer the boundary
This is where the current debate keeps failing. Pause or accelerate. Open or closed. Regulate the labs or trust the market. Pick a tribe and spend the next six months shouting at the other one. None of those binaries describe the actual system.
The governing boundary has to follow capability into consequence. The answer is not to centralize advanced AI behind three corporate gates. That gives private institutions too much authority and creates exactly the sovereignty problem Zuckerberg is warning about. The answer is also not to treat distribution as self-regulating. Once capability becomes portable, governance has to survive the loss of the provider.
That requires controls attached to the places where authority changes hands. I would start with capability gates. A developer releasing a frontier model owns the capability evaluation before release. The timebox is the release decision itself. An independent evaluator or designated public authority witnesses the result. The evidence is a retained, signed evaluation record tied to the exact model being released. A declaration that the model is safe is not evidence.
Agent deployment needs its own gate. The organization giving an agent access owns the permission boundary. Before deployment, tools, credentials, network egress, installed skills, persistence, and revocation paths are enumerated and tested. A security or platform owner witnesses the gate. The evidence is the permission manifest, provenance record, and test result. If nobody can show what the agent was allowed to do, there was no control.
Biological synthesis needs a physical gate. The synthesis provider and laboratory own the transition from generated sequence to manufactured material. Screening and research approval happen before synthesis, not after an interesting result appears in a paper. The witness is the accountable biosafety or biosecurity function. The evidence is the screening record, approval record, and chain of custody.
Boundary failures need an incident gate. I would require operators to preserve evidence immediately, make an initial report to the relevant public authority within 24 hours for high-consequence escapes or unintended external compromise, and produce the technical record within seven days. The operator owns the report. The regulator or designated authority witnesses it. Logs, model version, permissions, decisions, containment actions, and known impact become evidence rather than whatever story survives six months of legal review.
Those numbers are policy proposals, not existing universal requirements. The structure is the point. Owner. Timebox. Witness. Evidence. Without those four things, “AI safety” remains a collection of intentions.
The actual race
The arms race is not Zuckerberg versus Sanders, or even Meta versus OpenAI versus Anthropic. Zuckerberg’s own essay makes clear that frontier developers already think in terms of competitive advantage measured in weeks or months, national leadership, access to compute, and the risk of rivals moving faster. Sanders is responding by asking companies to stop and by trying to slow the infrastructure feeding that competition. (Meta)
But another race is already underneath that one. Capability is becoming easier to move while controls remain tied to institutions. Attackers and defenders are both getting better tools. Agents are receiving more authority. Open models can run outside provider controls. Digital instructions can cross into physical systems. Every one of those transitions creates another place where somebody has to own the consequence.
That is where governance has to live. Not in whether the technology feels frightening, whether the model is open or closed, or whether a CEO promises to stop if things get dangerous enough. The test is whether someone can name the boundary, the authority crossing it, the person responsible for that authority, the evidence that the control ran, and what happens when it fails.
If we cannot answer those questions, we are not governing the capability. We are watching it.
Sovereignty for users. Liability for operators.
A pause without an enforceable boundary is a request. Distribution without an enforceable boundary is an arms race. Neither is governance.
Artifacts are cheap, judgement is scarce.
Per ignem, veritas.



