NSA, CISA, FBI Name Six Chinese Firms Running Industrial AI Distillation
The advisory frames systematic extraction of proprietary model capabilities not as a cybersecurity incident category but as an economic-warfare supply-chain threat, a first for the three agencies acting jointly.
TL;DR
NSA, CISA, and the FBI identified six Chinese AI firms, DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI, conducting industrial-scale knowledge distillation of U.S. frontier AI models since at least late 2024. The advisory says distillation forms "the critical core" of these firms' development strategy, not a supplement to it. The companies used proxy "transfer stations," bulk-purchased premium subscriptions shared across developer teams, and automated failover between pathways to extract billions of tokens across millions of requests from Claude, GPT, Gemini, and Grok variants. The agencies recommend U.S. AI vendors implement anomalous-prompt detection, deploy targeted response alterations to degrade distillation payoffs, and establish cross-organization intelligence sharing.

The advisory names six firms and catalogs their targets in granular detail. DeepSeek, the most extensively documented, ran organized campaigns from late 2024 through mid-2025 against Claude 3.7, Sonnet 4, Sonnet 4.5, Opus 4.1, Gemini 2.5 Pro and Flash Preview, GPT-4, GPT-4o, GPT-4 Mini, GPT-4 Nano, GPT-5, and Grok 4. The extracted capabilities include legal specialization optimization, chain-of-thought reasoning extraction, agentic functions, and supervised fine-tuning optimization. DeepSeek's publicly quoted $5.6M training cost figure, the advisory notes, excludes the value of the data obtained through the distillation campaigns.
Moonshot AI extracted Claude Fable 5 data to train its Kimi-K3 model and GPT-4o data for Kimi-K2, targeting software engineering, math, reinforcement learning, and SFT optimization capabilities. Alibaba, MiniMax, StepFun, and Z.AI ran parallel campaigns. Alibaba distilled Claude-4, Opus, Sonnet, and GPT-5 to improve software engineering, customer service dialogue, and image generation. MiniMax extracted chain-of-thought reasoning and software engineering capabilities for its M2 model from Claude Code, Claude Sonnet 4, Claude Opus, and multiple Gemini versions, including using Claude Code for internal software development tasks.
The operational playbook
The advisory describes a layered evasion architecture. Chinese firms route distillation requests through native APIs, remote cloud providers, and third-party aggregators that automatically obfuscate user metadata. They use a gray market of proxies called "transfer stations" to bypass geographic restrictions. Bulk procurement of premium subscriptions (shared across teams of developers) reduces per-token costs. When one pathway is blocked, automated failover shifts traffic to another. The firms also deploy quality evaluation frameworks to detect when U.S. providers are deploying countermeasures that degrade extracted responses.
The agencies' recommended mitigations are specific: monitor subscription-to-usage ratios, flag immediate maximum usage from new accounts, identify enterprise-scale throughput patterns, and subtly alter responses for suspected malicious distillation attempts to attenuate the payoff. The third recommendation (cross-organization intelligence sharing across model providers, cloud platforms, and API aggregators) acknowledges that no single vendor can see the full campaign.
What's new here
This is the first joint NSA-CISA-FBI advisory to treat AI model extraction as a supply-chain threat rather than a routine cybersecurity incident. The framing matters: the advisory says the campaigns "likely with the knowledge of the Chinese government" and that the scale and sophistication are inconsistent with purely private-sector activity. The agencies are naming individual firms, specific models targeted, and the capabilities extracted, detail that would typically remain classified or in confidential industry notifications.
The advisory lands in a broader enforcement context. DOJ's Operation Gatekeeper seized over $50 million in advanced GPUs destined for China in December 2025. In March 2026, three defendants were charged with conspiring to divert AI-integrated servers to China. And NIST's CAISI evaluation of DeepSeek models in September 2025 flagged security shortcomings and censorship risks while noting the models' rapid global adoption. The distillation advisory connects the hardware-smuggling and model-evaluation threads: the goal isn't just acquiring chips or producing competitive models, it's extracting the proprietary capabilities embedded in U.S. frontier models at a scale that shortens adversary development timelines and erodes the technical moat.
For defense contractors and ISVs building on U.S. frontier models, the operational implication is that proprietary edge is eroding faster than public benchmarks suggest. The advisory doesn't impose mandatory mitigations (the agencies recommend, they don't require) but the specificity of the named firms, pathways, and techniques reads as a predicate for further action, whether entity-list designations or sanctions against the proxy infrastructure.
Published ·Deep Fathom