OpenAI has identified and disrupted a coordinated campaign designed to extract protected reasoning from its models, an activity the company attributes primarily to individuals associated with Moonshot AI, the developer of Kimi. The campaign, which began in early July, involved the systematic and unauthorized use of one model’s outputs to help train or improve another.

What Happened

The earliest observed activity occurred in the first week of July, starting at a low volume before spiking significantly on July 24 and 25. During this peak, OpenAI observed 16,000 requests using a relevant extraction pattern from over 4,000 users. Further investigation revealed related prompt-pattern activity across a cluster of more than 15,000 users, which was fully disrupted by July 28. The operators did not break encryption or compromise databases; instead, they manipulated model interactions to reproduce protected reasoning in visible forms, violating OpenAI’s terms of service. Specific techniques included copying encrypted reasoning from one conversation and asking a model in another conversation to decrypt and transcribe the hidden content.

Why It Matters

OpenAI frames adversarial distillation as a significant security challenge that poses safety and national security risks. When reasoning is extracted, it can reveal information withheld from final answers, allowing competitors to reproduce capabilities without necessarily preserving the safeguards applied to the original model’s user-facing outputs. This process can accelerate the transfer of advanced capabilities without requiring equivalent investment in safety, a concern that heightens as models gain capabilities in dual-use domains. The company notes that this risk is not unique to OpenAI, as similar techniques may affect other advanced AI systems, necessitating coordination across the industry to strengthen collective defenses.

The Bottom Line

OpenAI has implemented account enforcement, technical controls, and partner coordination to mitigate the campaign. The company expects adversarial distillation attempts to become more sophisticated and plans to continue improving tool defenses, classifier coverage, and threat-information sharing through the Frontier Model Forum and government channels.