ai-cybersecuritytrade-pressNewsThe Broadside2 min read

PLA cyber unit used GPT-3.5 to train classified military AI

Distillation extracts the reasoning without the chips, exposing a gap in the export-control framework built for hardware.


TL;DR

PLA Unit 96941 researchers used GPT-3.5 to summarize classified military code, then trained a domestic model on those summaries to run inside Chinese networks. The paper is one of 80-plus reviewed by Reuters and the Jamestown Foundation that show Chinese defense institutions systematically distilling U.S. AI models for surveillance, cyber ops, and tactical decision-making. The technique requires API access, not restricted chips, and the export-control regime wasn't built to stop it.

Reuters and the Jamestown Foundation reviewed more than 80 Chinese academic papers and patents and found defense-linked researchers systematically using model distillation to extract capabilities from U.S. AI systems. The most direct example: researchers in PLA Unit 96941, a Beijing-based military intelligence and cyber-warfare unit, used OpenAI's GPT-3.5 to process classified military source code. Because third-party models couldn't handle sensitive information on Chinese networks, they had GPT-3.5 summarize the code and trained a domestic model on those summaries, capturing the reasoning while keeping the system air-gapped.

The U.S. has spent years building an export-control architecture around chips. The October 2022 BIS rules restricted advanced semiconductor exports. The CHIPS Act poured $52 billion into domestic fabrication. A July 2026 executive order tightened defense supply-chain waiver rules for Chinese materials (defensenews.com). DOJ has prosecuted hardware smuggling: in December 2025, Operation Gatekeeper seized more than $50 million in Nvidia GPUs bound for China (justice.gov); in March 2026, three were indicted for conspiring to divert AI-integrated servers to Chinese buyers (justice.gov). But distillation doesn't need smuggled chips. It needs API access and patience. The models sit in U.S. data centers; the reasoning walks out through the front door.

The papers span more than military applications. Researchers at North University of China, which has close ties to the country's weapons industry, used Anthropic's Claude 3 Haiku to generate synthetic training data for social media monitoring and content moderation models. Other papers describe distillation for cyber operations and tactical decision-support systems. Sunny Cheung, the Jamestown fellow who analyzed over 60 of the papers, said the researchers aren't just extracting answers, they're capturing the reasoning steps, the expensive proprietary logic that distinguishes frontier models.

The dispute is already a flashpoint in U.S.-China AI governance talks. U.S. officials call it IP theft that undermines export controls; Beijing calls it AI "hegemonism." Meanwhile, the practice continues. Distillation is legal, widely used in industry, and nearly impossible to detect at the API level. Washington built walls around the hardware. The reasoning left through the API, and there's no gate.


Published ·Deep Fathom

PLA cyber unit used GPT-3.5 to train classified military AI — The Broadside