OpenAI Halts Extraction of Protected AI Reasoning by Moonshot AI-
AI-generated image

OpenAI Halts Extraction of Protected AI Reasoning by Moonshot AI-

OpenAI disrupted a coordinated campaign that attempted to extract protected reasoning from its AI models, attributing the activity to individuals tied to Beijin…

OpenAI has shut down a systematic campaign designed to illicitly extract protected reasoning from its artificial intelligence models, a technique known as adversarial distillation. The company linked a core part of the activity, spanning early July to late July 2026, to individuals associated with Moonshot AI, a Chinese AI firm based in Beijing.

The campaign began on July 1, 2026, with low-volume probe requests. It surged dramatically on July 24 and 25, when OpenAI recorded around 16,000 attempted requests using a specific extraction pattern spread across more than 4,000 user accounts. A wider sweep later identified related prompt-pattern activity across over 15,000 users. OpenAI fully disrupted the operation on July 28, 2026.

Adversarial distillation is the systematic and unauthorized use of one model’s outputs to train, replicate, or enhance another model. In this case, the operators did not break encryption, compromise databases, or gain direct access to stored conversations. Instead, they manipulated model interactions so that protected reasoning—the internal step-by-step thinking a model performs before delivering a final answer—could be reproduced in forms visible to the requester, all at scale and in violation of OpenAI’s terms of service.

To understand the risk, protected reasoning offers a window into how a model works through a task. Extracting it can reveal sensitive data and help competitors copy a model’s capabilities without investing the same resources in safety safeguards. OpenAI noted that extracted reasoning could be used to train another model while preserving the original’s performance but discarding the user-facing safety filters, which raises both security and national safety concerns, particularly as models gain capabilities in dual-use domains.

Following the incident, OpenAI deployed additional mitigations. It closed a “pathway” that had allowed someone who possessed another user’s encrypted reasoning trace to replay it and recover its plaintext contents. The company also added checks to detect and hold any streamed output that might expose reasoning before it reaches the user. Furthermore, OpenAI banned the fraudulent accounts involved in the campaign.

The episode highlights a deeper architectural issue. In a study published in August 2026, researchers from MATS Research, ELLIS Institute Tübingen, and Synk discovered a vulnerability affecting AI models from multiple providers, including Claude, Gemini, and GPT. They found that encrypted reasoning traces are fully compatible and interchangeable across different sessions, users, and models within a single provider’s ecosystem. An attacker could inject an encrypted reasoning trace from a strong, well-guarded model into a weaker, less safeguarded model from the same provider, forcing the weaker model to decode and output the reasoning verbatim in plaintext—without ever directly jailbreaking the stronger model. This technique also enables large-scale private data extraction, invisible prompt injections where malicious payloads are hidden entirely within encrypted blocks, and the accidental exposure of hazardous information from the reasoning process even when the model’s final visible output rejects a harmful request.

This is not the first time Moonshot AI has faced accusations of distillation. Last month, rival AI company Anthropic alleged that Moonshot AI was stealthily relaying customer requests from its own Kimi platform to Anthropic’s Claude model, then displaying Claude’s responses back to users. Anthropic further claimed that Moonshot AI retained a subset of those exchanges to train its own chain-of-thought model. That activity is tracked under the identifier GTG-16002.

For businesses and website operators who integrate AI-driven features into their online presence—such as chatbots, content generators, or customer support automation—the incident is a reminder that third-party AI services are only as secure as their most stressed safeguards. AEU-I, our security-first IT and infrastructure consultancy, helps organisations assess the risks tied to third-party tools, implement data handling policies, and monitor for abnormal access patterns that could signal unauthorized extraction of proprietary information.

How to Protect Yourself

  1. Read the privacy policies and data usage settings of any AI tool your website uses, such as a chatbot or content generator, and switch off data sharing for model training if the option exists.
  2. Avoid pasting sensitive business data, customer details, or proprietary code into public AI prompt interfaces.
  3. Rotate API keys for AI services regularly and revoke any you no longer use to prevent unauthorized access.
  4. Monitor the usage dashboard of your connected AI accounts for unexpected spikes in requests, as those can indicate abuse of your account.
  5. Keep your website’s plugins and themes up to date so attackers cannot inject malicious code that intercepts interactions between your visitors and embedded AI features.

Terms Explained

  • distillation A technique where knowledge from a large, complex AI model is transferred to a smaller or different model, often by training it on the outputs of the larger model.
  • adversarial distillation The unauthorized use of one model’s outputs to replicate or improve another model, typically bypassing safety filters and terms of service.
  • encrypted reasoning The internal, step-by-step thought process an AI model goes through before giving a final answer, which is often hidden and protected to guard the model’s inner workings.
  • prompt injection An attack where someone hides a malicious instruction inside the input to a model, trying to make it behave in unintended ways or reveal hidden data.
  • chain-of-thought (CoT) A method where the AI articulates intermediate reasoning steps to improve its final answer, making its thinking more transparent.

Related AEU services