ai-cybersecuritytrade-pressNewsThe Broadside1 min read

OpenAI disrupts Moonshot-linked decrypt-and-distill campaign

The exploit turned encrypted reasoning into readable output across conversations, but OpenAI's Moonshot AI attribution comes with no public technical evidence.


TL;DR

OpenAI says it disrupted a coordinated campaign that copied encrypted reasoning data from one conversation and asked the model in a separate conversation to decrypt it, surfacing protected output without breaking encryption. It ties a "core cluster" to Moonshot AI but publishes no technical evidence. OpenAI fixed the cross-conversation bug, shared the finding with the Frontier Model Forum, and says other models are also vulnerable.

OpenAI's Wednesday disclosure lands three weeks after the joint CISA-NSA-FBI advisory named Moonshot AI among the Chinese companies running industrial-scale distillation campaigns against U.S. models. OpenAI now says it disrupted a coordinated campaign that ran from July 1 to July 28, with a "core cluster" of operators working on behalf of Moonshot AI.

The attribution is the soft part. OpenAI's unsigned blog post offers no technical evidence or reasoning, and it concedes the activity may not all be related. The method, on the other hand, is worth attention on its own. Operators copied encrypted reasoning data from one conversation, then asked the model in a separate conversation to decrypt and transcribe it. Nobody broke encryption or compromised a database. They manipulated model interactions so protected reasoning became visible to the requester, at scale, in violation of the terms of service.

OpenAI says it fixed the cross-conversation decryption bug, banned offending accounts, tightened signup and infrastructure controls, and expanded network monitoring. It also says the same vulnerability exists in other models and that it shared details with the Frontier Model Forum. For organizations tracking the CISA advisory's mitigation guidance, that's the operational signal: this was a prompt-level interaction exploit, and the advisory already called for detecting anomalous prompts and accounts before abuse scales.


Published ·Deep Fathom