Theft
Never before has so much power been concentrated in so little data.
On this page
I hunted out the source of fire, and stole it.
— Prometheus
GPT-6 Astra is only a few terabytes. Anthropic’s Mythos is roughly the same. We’re spending hundreds of billions on compute, electricity, and data to produce a simple set of files: the model weights. But these files can be stolen, and for the world’s most sophisticated intelligence services, stealing the weights offers a shortcut to superintelligence. America’s lead in the AI race could disappear overnight if someone steals its most advanced frontier model.
We must build nation-state-grade security for AI model weights to withstand the world’s most sophisticated AI-augmented intelligence services.
Why Steal a Model?
Frontier AI models are becoming the most valuable artifacts humanity has ever created. They are accelerating scientific discovery, solving mathematical problems that have remained open for decades, and powering software development. They could dramatically increase a nation’s productivity, GDP, and economic power. At the same time, the internal models at the frontier labs can provide instructions for building chemical, biological, radiological, nuclear, and explosive (CBRNE) weapons and power massive cyberswarms. All of this scientific, economic, and military power is compressed into a few terabytes that can be copied across a network. Never before has so much power been concentrated in so little data.
Because so much value is concentrated in such small files, model weights are becoming some of the most sought-after targets for cyber espionage.
Why would someone steal a model rather than develop it themselves? It allows them to gain two major advantages:
- Bypass the CapEx. A threat actor that steals model weights acquires the product of hundreds of billions of dollars in infrastructure, vast quantities of data, and years of research without reproducing the training process. Russia may not have the compute necessary to train a frontier AI model, but it certainly has the compute to run a stolen one. Once it possesses the weights, it can remove the model’s safeguards and use it for autonomous hacking, cyberwarfare, scientific research, weapons development, or the creation of more efficient AI systems that can run on more limited infrastructure. The original developer loses all control over how the model is modified, copied, or deployed.
- Steal the data inside the model. Models can reveal information from the data on which they were trained. In 2025, researchers used Llama 3.1 to reconstruct the first book in the Harry Potter series nearly verbatim. Now consider a model fine-tuned on classified intelligence, military data, patient records, or proprietary corporate information. Stealing its weights could also expose some of the sensitive data encoded within them.
Model-weight theft may be the highest-value act of intellectual-property theft ever made possible.
The primary state threat actor for model-weight theft is China. China has spent decades using insiders and industrial espionage to steal American technology. A former Google engineer was convicted of stealing thousands of pages of confidential AI technology for the benefit of China. In a more elaborate operation, a GE engineer concealed stolen turbine designs inside an ordinary image of a sunset and emailed it to himself.
How could a threat actor steal a model?
How to Steal a Model
For obvious reasons, we will not provide a blueprint for stealing a frontier model. But the broad attack paths are already public. An adversary has two options: take the model itself by exfiltrating its weights, or copy its capabilities through distillation.
Exfiltrating the weights
Frontier model weights are distributed across thousands of nodes inside an AI data center. If an attacker compromises one of those nodes and uses that access to copy the weights onto an external drive or server, it is game over.
That access could come through a cyber intrusion, a privileged insider, or both. An insider may know where the weights are stored, which systems can reach them, and how those systems are monitored. They could steal the weights directly or create access for an external cyber team. A nation-state operation could combine both: the insider opens the door, and the cyber team moves the model.
Social engineering can turn trusted employees into unwitting insiders. Attackers may appear to be “friends” with people from the labs and, in the process, gain a great deal of information about daily processes, the types of security in the building, the limitations in that security, and which buildings or locations the models are stored in. They may also listen to conversations held on a bus, in a restaurant, or in a coffee shop. This is a genuine point of concern for us. Frontier lab researchers are the custodians of strategic assets, and they need counterintelligence training: how to detect social engineering attempts, and how to safeguard the operations and strategy of their companies.
Simply sending several terabytes out of the data center should be detectable. When Anthropic released Claude Opus 4, one of the more than 100 security measures it activated was a preliminary system of egress-bandwidth controls, which limit and monitor data leaving systems that contain model weights. Because a full frontier model is several terabytes, these controls can flag or block the sustained transfer before it finishes.
An attacker could try to evade egress controls by slow-dripping the weights over days or weeks, breaking them into small fragments, and hiding those fragments inside routine outbound traffic such as API calls or telemetry sent to services like Datadog. The data could also be distributed across multiple destinations, making each individual transfer less conspicuous. Alternatively, a compromised data-center employee — or anyone with sufficient physical access — could bypass network monitoring entirely by copying the weights onto physical storage and carrying them out.
There are also more advanced covert methods. One publicly discussed example is steganography: altering the way tokens are sampled from a model so that hidden information is encoded inside outputs that appear normal. In principle, those outputs could be used to leak information about the model — including portions of its weights — without producing an obvious large-file transfer. Researchers have also covertly transmitted model weights through ordinary wireless traffic and reconstructed the model externally. There are further, more advanced techniques that we will not discuss here.
Copying the capabilities: distillation
The second path does not require breaking into anything. Model distillation is when you copy a model by training on its outputs. You are not copying the model file itself, but rather copying the model’s capabilities and trying to insert them into your own model. It has been proven on many occasions that China is distilling American models. To be clear, we do not believe this is the main reason China is succeeding in increasing its model capabilities — China has a very strong team of R&D talent, and a much better energy and grid situation than the United States. But distillation is prevalent today, it is cheap, and unlike weight theft it requires no intrusion at all: just querying an American model through its ordinary API and keeping the data. We return to distillation, and to what the United States might do about it, in Offense.
Weights have already walked out the door
Model weights have already escaped through trusted access. In March 2023, Meta’s LLaMA weights leaked online within days of being distributed to approved researchers. Meta provided the weights on a case-by-case basis, and an unknown person posted a torrent containing them on 4chan on March 3, which rapidly spread through AI communities. The leaker was never publicly identified. (The Verge, the grugq) In January 2024, Mistral’s CEO confirmed that an employee at one of its early-access customers leaked a quantized, watermarked version of an older model, which was posted to Hugging Face and leaked on 4chan. (VentureBeat, Analytics Vidhya) Neither required penetrating a hardened cluster — just authorized access. Both confirmed weight leaks followed the insider-with-legitimate-access pattern: the failure was access control and distribution discipline, not external intrusion or state-sponsored espionage.
The surrounding systems have also been penetrated. In 2023, a hacker stole technical discussions from OpenAI’s internal forum. In 2024, the ShinyHunters-associated UNC5537 campaign used stolen credentials to exfiltrate data from roughly 165 Snowflake customer environments, while Wiz demonstrated a cross-tenant vulnerability at Replicate that could have exposed customers’ prompts, outputs, and private models. In 2025, a former xAI engineer acknowledged copying confidential files, including the company’s source-code repository, and a former Google engineer was convicted of economic espionage for stealing more than 2,000 pages of AI-supercomputer trade secrets for the benefit of the Chinese government.
A hostile actor has not yet been publicly confirmed stealing a closed frontier model directly from a lab. But insiders have taken the knowledge used to build one, hackers have breached the systems surrounding one, and researchers have demonstrated how its weights could leave undetected.
The danger is magnified by the speed at which AI infrastructure is being deployed. Demand for compute is effectively infinite, and companies are building clusters as quickly as they can acquire chips, energy, land, and networking equipment. At the moment, these systems are not even close to being secured against sophisticated model exfiltration by nation-state actors.
Put bluntly: if the full force of China’s Ministry of State Security, China’s civilian intelligence service, were directed at stealing the most sophisticated AI model in the world, likely from OpenAI or Anthropic, it could probably obtain the model through an insider threat, purely cyber means, or a combination of both.
As the previous sections show, these threats overlap. Sabotage requires enough access to read, modify, or replace the weights, but it does not require moving them outside the lab. In model theft, gaining that access is only the first half of the attack; the attacker must also copy the weights and get them out. Escape follows the same path, except the model itself is the attacker: once it gains command-execution access on an external machine, it can use that machine to conduct the same kind of intrusion and exfiltration as a human attacker. Sabotage, escape, and theft therefore converge on the same critical security boundary: who or what can access the weights, and whether it can move them.
The Worst-Case Theft: The Theft of RSI
Consider the following highly possible scenario:
-
October 2026
An adversarial intelligence service establishes a persistent exfiltration path inside the infrastructure of a frontier U.S. AI lab. It doesn’t take anything, so no weights leave the lab. Doing so would reveal its access, trigger an investigation, and cause the vulnerability to be closed. It waits.
-
March 2027
The lab trains the first model that crosses a critical recursive self-improvement (RSI) threshold. RSI begins when an AI model can conduct AI research and create an even smarter AI model. That smarter model can then create an even smarter model, which can create an even smarter model. Intelligent machines create smarter machines, which create smarter machines. The result is an intelligence explosion.
-
That same day
The persistent exfiltration path is activated. The weights are copied and transferred before the lab realizes that its infrastructure has been compromised. The adversary now possesses the same model.
This is the worst-case scenario of theft. Why? If the Chinese Communist Party, or any adversary, steals the weights for an RSI-capable model, it could use that model to leapfrog in AI capability, rapidly close the gap with the United States, and potentially pull ahead. Any American lead in AI could disappear overnight.
China is the primary risk. It already possesses a mature energy ecosystem and generates roughly 2 to 2.5 times as much electricity as the United States. While energy is becoming one of the primary constraints on AI infrastructure build-out in the United States, China is building it at an extraordinary scale. Advanced chips remain export-controlled to China, but it is developing a domestic chip pipeline through firms such as Huawei. By the time RSI is created, that pipeline may be mature enough to reach the FLOPS-per-watt threshold and throughput required to train or fine-tune stolen models at scale.
Moreover, an RSI-capable model may not necessarily require overwhelming compute. If it can make major algorithmic improvements while using less compute, an adversary could potentially leapfrog the United States despite possessing fewer advanced chips.
The event that causes the United States to lose the AI race may not be China training a better model. It may be China stealing the weights for America’s first RSI-capable model.
Call to action 3
Nation-state-grade security for model weights
We must build nation-state-grade security for AI data centers. These systems must protect model weights against cyberattack, insider threats, covert model theft, and operations that combine all three.
We need advanced cyber defenses for AI data centers, and we need new technologies that keep model weights secure at every stage of model development and deployment. The weights must remain protected even if an individual machine, employee, or infrastructure layer is compromised.
We call for an urgent national mission to build advanced defenses for AI data centers and model weights, aimed at preventing the theft of advanced AI models and stopping these capabilities from leaking to our greatest geopolitical adversaries.
AI will be the defining defense capability of the century. It will redefine cyberwarfare, biological and chemical warfare, intelligence operations, and kinetic warfare.
We must protect the weights.