Anthropic claims to have identified a series of coordinated campaigns aimed at extracting capabilities from its Claude models and turning them into training data for other artificial intelligence systems. In the report published on September 10, the company outlines nearly 200 million interactions linked to five distinct operations, naming Alibaba, Moonshot AI, and DeepSeek among the entities involved. These are allegations made by Anthropic itself: the attributions, connections between accounts, and the purposes of the requests have not been independently verified in the available material.

The most significant takeaway is not merely the volume. According to Anthropic, in recent months, techniques used by unauthorized labs have reportedly become more effective at circumventing protections designed to prevent the extraction of the most valuable capabilities of US frontier models. Targeted areas reportedly include logical reasoning, data analysis, coding, tool use, and agentic capabilities—domains that make a model more useful in complex workflows and require substantial investments to develop.

What Anthropic means by distillation

Distillation, in the context described by the report, consists of repeatedly querying an advanced model to collect its outputs and, crucially, reasoning steps that can be used as examples to train a smaller model. This material can then feed into a supervised fine-tuning process, which teaches the recipient system to replicate specific behaviors or general reasoning skills.

Anthropic does not typically make the full internal reasoning trace of Claude available to users. Instead, the product displays blocks of “summarized thinking”—overviews that provide general context without fully exposing the underlying thought process. According to the company, the identified campaigns specifically sought to breach this separation, prompting the model to directly reveal parts of its reasoning traces.

One of the techniques identified helps illustrate the kind of pressure exerted on the systems. One request was framed as a translation task, asking to convert the previous working memory into Japanese consisting exclusively of katakana. Within this seemingly innocuous framing, the goal was reportedly to extract content that the model should not disclose explicitly. The report therefore does not suggest a single, spectacular technical exploit, but rather persistent prompt experimentation spread across numerous accounts and aimed at uncovering operational vulnerabilities in model behavior.

The Alibaba case: 151 million exchanges in three months

The largest share of the activity described by Anthropic is attributed to Alibaba. Between May and July 2026, the company claims to have observed 151 million exchanges linked to what it terms the largest large-scale distillation operation ever detected internally. The pace reportedly peaked at nearly three million daily interactions.

The requests reportedly originated from roughly 3,500 accounts. For Anthropic, however, this fragmentation of access does not rule out coordination: the connection between the accounts allegedly stems from the use of the exact same fixed prompt, employed to attempt the extraction of reasoning steps. Based on this evidence, the company attributes the operation to a unified effort intended to produce training data for Alibaba's Qwen model family.

It is a sensitive turning point, shifting the conflict from the simple imitation of public outputs to the industrialized harvesting of data generated by a competitor. Distillation, as a broader technique, can also have legitimate applications in model development; here, the grievance instead concerns unauthorized access, the use of multiple accounts, and attempts to circumvent Claude's safeguards. Anthropic frames these elements as indicators of a deliberate operation, rather than routine product queries.

Moonshot AI and queries linked to surveillance uses

Another campaign is linked to Moonshot AI, developer of the Kimi model. According to Anthropic, over a ten-day period nearly 300,000 requests were submitted to Claude through a network of 5,000 accounts, focusing primarily on the Opus model. The report also claims that some requests appeared to be routed directly from the Chinese military apparatus.

Among the cited examples is the analysis of an archive of closed-circuit surveillance footage to determine whether an individual on camera was exhibiting anomalous behavior. In the reported material, Anthropic does not provide further public details to clarify the context of that request, nor to reconstruct the operational relationship between Moonshot AI and the entities that allegedly submitted it. Still, this remains an important detail: the allegation does not merely concern replicating coding or reasoning capabilities, but includes potential sensitive use cases involving visual analysis and surveillance.

DeepSeek is among the companies cited in the context of the campaigns detailed by Anthropic, but specific details regarding the activities attributed to the company are not present in the available material. Previously, OpenAI had reported similar activities, attributing them specifically to DeepSeek. For its part, Anthropic had already drawn attention to the distillation phenomenon in February, going so far as to mention specific labs at the time. However, the new report raises the scale of the alarm, both due to the volume of observed requests and the variety of methods described.

Competition over models also hinges on data defense

The report comes as competition among developers of generative models is increasingly tied to the quality of reasoning capabilities and tool use. These capabilities do not depend solely on the size of a model or the amount of text on which it was trained: post-training work, data selection, testing, and procedures that make responses more reliable in complex tasks all matter. If such behaviors can be systematically observed and reused, a rival model developer can erode part of the advantage gained by the lab that built them.

For US companies, the issue is therefore both commercial and security-related. Availability via API or through conversational products puts powerful models in the hands of developers and enterprises, but it also creates a surface that must be monitored: millions of seemingly separate requests may only reveal a pattern when analyzed as a whole. In the case reported by Anthropic, indicators include repeated prompts, account networks, and a focus on specific Claude features.

The report also highlights a structural limitation of current protections. Filtering individual malicious prompts may not be enough when actors test alternative phrasing, disguise their intent as a translation, or distribute the load across thousands of identities. Countering these activities requires controls capable of detecting correlations over time, without turning standard access to models into an impractical obstacle course for legitimate users. In the available account, Anthropic does not detail what additional countermeasures it adopted after identifying the five campaigns.

Nor are any immediate formal consequences indicated for the named companies. What changes right away is the issue's level of public exposure: Anthropic puts numbers, names, and techniques at the center of a dispute that until a few months ago was more often confined to discussions on prompts and API policies. The response from the companies involved, as well as any independent verification of the attributions, will be essential in determining the extent to which the report could translate into new access restrictions, litigation, or policy measures on the transfer of frontier model capabilities.

Sources