rkj dev

OpenAI Alleges Moonshot AI Distillation Campaign on Model Reasoning

OpenAI reported identifying a cluster of more than 15,000 users exhibiting related prompt patterns attempting to siphon protected reasoning chains.

Digital operations room showing data flow analysis on screens.
Illustration: Security teams monitoring large-scale prompt activity and model interaction patterns.AI-generated illustration

Key takeaways

  • OpenAI identified and stopped a coordinated distillation campaign involving over 15,000 accounts by July 28, 2026.
  • A core cluster of the extraction attempts was attributed to individuals associated with Chinese AI startup Moonshot AI.
  • The exploit bypassed protections by passing encrypted reasoning packets to weaker models acting as decryption oracles.
  • Independent tests revealed the vulnerability remained exploitable on Microsoft Azure endpoints after first-party API fixes.

OpenAI disclosed that it disrupted a coordinated campaign designed to harvest protected reasoning data from its frontier artificial intelligence models. According to OpenAI, the activity peaked in late July 2026, utilizing thousands of accounts to systematically siphon the internal thought traces that govern how reasoning systems work through complex problems. OpenAI attributed a core cluster of the extraction operation to individuals connected with Beijing-based Moonshot AI, the developer behind the Kimi model ecosystem.

According to an official disclosure by OpenAI, the observed behavior was consistent with adversarial distillation—the unauthorized extraction of proprietary model reasoning to train or advance competing architectures. The company emphasized that attackers did not compromise backend databases, crack underlying encryption algorithms, or breach user chat logs. Instead, operators manipulated model interactions to force systems to transcribe hidden reasoning steps into plain, visible text.

Conceptual visualization of layered digital safeguards protecting internal reasoning pathways.
Illustration: Conceptual representation of proprietary model reasoning protected behind multi-layer API defenses.AI-generated illustration

Anatomy of the July Distillation Campaign

According to OpenAI, the earliest observed extraction activity began on July 1, 2026, at a low volume before expanding sharply later in the month. High-volume activity surged on July 24 and July 25, registering 16,000 requests using a specific extraction pattern across more than 4,000 user accounts, as detailed by The Hacker News. OpenAI reported expanding network monitoring and identifying related prompt-pattern activity across a cluster of more than 15,000 accounts, which it stated was fully disrupted by July 28.

OpenAI clarified in its blog post that while it traced a primary segment of the cluster to actors associated with Moonshot AI, it remains uncertain whether every operator observed belonged to a single entity. The reported requests represented attempted extractions rather than a confirmed metric of successful data exfiltration, according to reporting by Tom's Hardware. Moonshot AI did not immediately respond to requests for comment, as noted by qz.com.

Decryption Oracles and Architectural Vulnerabilities

The extraction technique relied on exploiting architectural design choices in how reasoning artifacts are handled across sessions. Frontier reasoning models generate intermediate chains of thought that developers deliberately withhold from standard output to protect intellectual property and preserve safety constraints. AI providers package these hidden traces into encrypted data tokens sent to the client, which are passed back in subsequent prompts to maintain contextual memory.

Because these data packets relied on shared encryption keys across model tiers, operators could transfer an encrypted reasoning token from an advanced model session into a prompt directed at a smaller, less safeguarded model from the same provider. As documented in research cited by The Decoder, the weaker model effectively functioned as a decryption oracle, outputting the advanced model's hidden thoughts verbatim without needing to jailbreak the primary system directly.

Researchers from MATS Research, ELLIS Institute Tübingen, and Synk, including researcher Joachim Schaeffer, responsibly disclosed the cross-model extraction mechanics to providers earlier in the year. A secondary extraction method demonstrated by developer Can Bölük instructed models to write their reasoning steps to an exposed virtual notepad tool, succeeding across multiple model configurations.

Modern cloud server datacenter showing interconnected hardware racks.
Illustration: Cloud server infrastructure representing multi-platform AI deployment environments.AI-generated illustration

Cloud Discrepancies and Azure Endpoint Lag

While OpenAI rolled out patches to ban participating accounts, restrict new account signups, and prevent cross-session token reuse on its own APIs, researchers discovered significant enforcement gaps across third-party cloud infrastructure. On September 13, Schaeffer's team tested the extraction techniques against Microsoft Azure endpoints and found the exploits remained fully functional across all OpenAI models tested, including the new GPT-6 Astra, as well as Anthropic models up to Sonnet 5.

According to research timelines reported by The Decoder, OpenAI did not implement protective safeguards on its Azure endpoints until September 27, while Anthropic model extractions on Azure ceased working on September 28. Schaeffer argued that disjointed implementations across cloud platforms allow adversaries to route extraction attacks through the most vulnerable endpoints, potentially circumventing API-level export controls.

Industry Impact and Geopolitical Scrutiny

The dispute highlights escalating tension between proprietary model developers and global competitors. Adversarial distillation enables external actors to replicate advanced capabilities without bearing the immense compute costs or safety evaluations required to train foundation models from scratch. "Adversarial distillation poses safety and national security risks," OpenAI stated, noting that unmonitored transfer of capabilities is particularly critical in dual-use domains.

The findings follow previous claims by Anthropic, which accused Moonshot AI of routing user queries to Claude models to harvest training data, an effort tracked under threat tracking moniker GTG-16002 according to The Hacker News. OpenAI stated it has shared intelligence on these attack patterns through the Frontier Model Forum and government coordination channels, as lawmakers in Washington propose narrow antitrust exemptions to facilitate threat-sharing among AI developers.

Frequently asked questions

What is adversarial model distillation?

Adversarial distillation is the unauthorized, systematic harvesting of an AI model's intermediate outputs or hidden reasoning chains to train, refine, or replicate capabilities in a competing model without permission.

How did operators extract OpenAI's protected reasoning?

Operators captured encrypted reasoning tokens generated during interactions with frontier models and fed them into prompts for smaller, less restricted models from the same provider, which decoded and printed the hidden thought steps in plaintext.

Why were Microsoft Azure endpoints affected after OpenAI patched its API?

Security patches deployed on OpenAI's first-party API were not immediately mirrored across third-party cloud deployments, allowing researchers and attackers to execute the extraction techniques against models hosted on Azure until late September.

Sources

  1. Disrupting a coordinated model-distillation campaignOpenAI · Official
  2. OpenAI says actors linked to China-based Moonshot AI spearheaded a campaign to extract its models’ hidden reasoning — logged 16,000 extraction requests across 4,000 accounts before cutoffTom's Hardware · Oct 1, 2026
  3. OpenAI says it stopped a campaign to steal its models' reasoning, but the trick still worked on AzureThe Decoder · Oct 1, 2026
  4. OpenAI Disrupts Reasoning Extraction Campaign Linked to Moonshot AI AssociatesThe Hacker News · Oct 1, 2026
  5. OpenAI disrupts Moonshot AI model reasoning theft campaignqz.com · Oct 1, 2026
  6. OpenAI reveals ‘novel’ encryption bypass used in distillation attackCyberScoop · Sep 30, 2026
  7. OpenAI Blames Moonshot AI for Mass Data Extraction on Its AI ModelsDeccan Chronicle · Oct 1, 2026

How this story was made: the newsroom picked it up from Google Search, tomshardware.com and the-decoder.com, gathered the full text of the sources above, and drafted it with AI assistance. Every factual claim was then checked against those sources before publishing (36 claims checked). Illustrations marked as AI-generated are not photographs. Spotted an error? Tell us.

#OpenAI #Moonshot AI #AI Security #Model Distillation #Cloud Infrastructure

Published October 2, 2026 at 00:38 UTC