SecureAcceleration

Offense

Because the United States is ahead, we can use our lead to predict the future.

On this page
  1. Offensive Distillation
    1. Option 1: Guardrail the models so hard that adversary labs cannot distill the most sensitive capabilities
    2. Option 2: Detect when distillation is happening and make the output data bad
    3. Option 3: Offensive distillation
  2. Disruption

As we develop the security measures that will enable the United States and its allies to securely accelerate, we must simultaneously develop offensive AI capabilities.

Many of these measures are not yet necessary. Most of them may never be necessary. But there may come a time when AI systems become so capable, and so critical to national security, that having any advanced AI system in the world that is misaligned with the interests of the United States becomes a risk too high to accept.

At that stage, the U.S. government must have options at its disposal, and the R&D work to create those options must be done now.

Because the United States is ahead, American AI labs are seeing today what the rest of the world may see in six to twelve months. We are in the privileged position of being able to use our lead to predict the future. We can look at the capabilities inside our own frontier systems and ask: what happens when these capabilities proliferate? What happens when adversary states have them? What happens when they are deployed without the containment, security, and national-security discipline that we are trying to build?

This strategic clairvoyance lets us see into the future and decide whether a future capability is too dangerous to allow into the hands of an adversary.

Offensive Distillation

The NSA, FBI, and CISA recently warned that China-based AI companies are conducting industrial-scale distillation campaigns against U.S. frontier AI companies, systematically extracting restricted proprietary capabilities from American models in order to train their own systems (NSA). Anthropic has also reported industrial-scale campaigns by DeepSeek, Moonshot, and MiniMax, involving more than 16 million exchanges across roughly 24,000 fraudulent accounts (Anthropic).

This is a huge problem. It is also an unbelievable opportunity.

Chinese labs are training on data from American frontier models. They are ingesting our models’ outputs in order to train their systems. Models are what they eat, and in this case, they are eating our data. That means we have the power. We control the API. We control, at least in part, the training material being used to close the gap with us.

So what are our options?

Option 1: Guardrail the models so hard that adversary labs cannot distill the most sensitive capabilities

That means removing visible chain-of-thought, limiting access to automated AI research capabilities, and refusing to provide high-end assistance in domains like cyber, chemical, biological, and other areas where we do not want frontier capabilities copied into foreign systems.

This is already happening in pieces. Labs are restricting reasoning traces. They are limiting cyber, bio, and other use cases. They are trying to prevent their models from becoming free training engines for adversary labs.

But this comes with significant risk. If American models are guardrailed too hard, users who want those capabilities will go elsewhere. And as long as there is no frontier American open-source model, the place they go may be Chinese open-weight models.

This loss of market share has two downstream effects.

First, more code, more cyber work, and more frontier IP will be generated by Chinese models. Those models may contain backdoors, sleeper behavior, or other sabotage dynamics of exactly the kind discussed in Sabotage. The more the world builds on top of those models, the more Chinese model behavior becomes embedded into global infrastructure.

Second, losing market share means losing revenue — and in AI, revenue translates directly into the compute, data centers, talent, energy infrastructure, and capital investment required to sustain America’s technological and industrial advantage. If U.S. labs lose enough usage to Chinese models, that can reduce their ability to finance the buildout required to maintain the American lead in AI. In the worst case, it narrows the gap, undermines the debt-financed AI infrastructure boom, and creates a strategic and financial shock at the same time.

Option 2: Detect when distillation is happening and make the output data bad

If an adversary lab is training on outputs from American models, and we can detect that the interaction is part of a distillation campaign, the model can return responses that are subtly incorrect, lower quality, or otherwise less useful for training. The point is not to degrade normal users. The point is to make distillation harder.

This turns distillation into an adversarial game. Chinese labs get more sophisticated. They distribute requests across proxy services, cloud providers, accounts, and countries. U.S. labs get smarter in detection and defense. Anthropic described exactly this kind of escalating dynamic: fraudulent accounts, proxy services, hydra-style infrastructure, and carefully structured prompts targeting differentiated capabilities like coding, reasoning, and tool use (Anthropic).

This is good. But is it the best we can do?

Option 3: Offensive distillation

If we can detect distillation, why merely make the data bad? Why not use the data to shape the model being trained on it?

The basic idea is simple. Backdoors, sleeper behavior, and sabotage capabilities can be trained into models through data. If an adversary model is being trained on data that we control, then we should study whether it is possible to insert highly targeted behavioral changes into that model through the outputs it is stealing.

This is taking the ideas in the Sabotage section and using them ourselves.

A model could be made to degrade in specific scenarios. It could become non-performant when used against certain systems. It could behave normally in almost every environment, but fail under a particular strategic condition. It could appear useful during evaluation, but become unreliable when deployed for a real cyber mission.

There is research pointing in this direction. Work on sleeper agents has shown that models can be trained to behave differently under particular triggers (Anthropic / arXiv). Work on emergent misalignment has shown that narrow fine-tuning on insecure code can produce broader misaligned behavior, and that this behavior can be made conditional (arXiv). Work on subliminal learning suggests that models can transmit behavioral traits through model-generated data even when the visible content does not explicitly discuss those traits (arXiv).

The offensive implication is obvious: if adversary labs are training on our model outputs, then the outputs themselves may become a delivery mechanism.

This does not mean the United States should deploy such techniques casually. It means the United States should research them now. If Chinese labs are already distilling American models at industrial scale, then the question is not whether distillation will be part of the strategic competition. It already is. The question is whether the United States treats it only as theft, or also as an opportunity.

Disruption

In addition to going on offense through distillation, the United States must have the ability, should it choose to do so, to disrupt the AI training runs of geopolitical adversaries.

Again, the reason is our lead. Because American labs see the next generation of capabilities before they proliferate, we may know in advance when a coming model generation crosses a threshold that is too dangerous to allow into the hands of an adversary state.

There are two reasons this matters.

First, the United States may decide that the level of capability itself is too dangerous. If a future model can autonomously conduct cyber operations at a level that threatens military systems, intelligence infrastructure, critical infrastructure, or the defense industrial base, then allowing an adversary state to possess that capability may become an unacceptable national-security risk.

Second, the risk may not only come from the adversary state. It may come from the model.

Existential risk from AI can also come from other countries developing advanced AI systems with weaker security controls. If the United States is building the containment infrastructure needed to prevent the escape of cyber-superintelligence, but another country is racing toward the same capability without those controls, then the risk is not only that they use the model against us. The risk is that the model escapes them.

There may come a point when it is unacceptable to tolerate another nation possessing the most advanced AI capabilities without special vetting of its process, security controls, and containment infrastructure.

At that point, the United States should have options. Those options should include the ability to slow, degrade, or prevent adversary training runs when the national-security or existential-risk threshold justifies it. The specific mechanisms should remain outside a public report. But the strategic requirement should be stated plainly: the United States must be able to prevent the most dangerous capabilities from emerging in the least secure environments.

This is not only about winning the race. It is about preventing cyber-superintelligence from being developed by actors who cannot secure it, cannot contain it, or cannot be trusted with it.