Anthropic Exposes Alleged AI Distillation Attacks by Alibaba, Moonshot AI and DeepSeek

TL;DR
- Anthropic said in a report released Thursday, September 10 that labs linked to Alibaba, Moonshot AI and DeepSeek ran persistent, large-scale distillation campaigns to extract capabilities from its Claude models.
- The alleged operations used fleets of fake accounts, evasive infrastructure and jailbreak prompts to harvest reasoning outputs for training rival models, with Anthropic saying it banned thousands of accounts in response.
- The accusations mark a sharp escalation in U.S.-China AI rivalry, raising new questions about model security, terms-of-service enforcement, and whether distillation will trigger tighter API controls and geopolitical fallout.
Anthropic Drops Its Most Direct Accusation Yet
In a detailed threat intelligence report published Thursday, Anthropic accused three of China's most prominent AI players — Alibaba, Moonshot AI and DeepSeek — of carrying out sustained campaigns to distill its Claude models.
The report is the company's most aggressive public attribution to date. Anthropic said the activity was not casual testing or academic evaluation, but industrial-scale extraction operations designed to fast-track competitor models by copying Claude's performance, reasoning and safety-evasion behavior.
All three named companies have not yet issued detailed rebuttals as of Friday, but the claims are already reverberating across Silicon Valley and Beijing amid intensifying competition over frontier models.
What Is Model Distillation — And Why It Matters
In AI development, distillation typically refers to the legitimate practice of using outputs from a larger, more capable model to train a smaller, more efficient one. Done internally, it is standard practice.
What Anthropic describes is unauthorized, cross-lab distillation at scale: using another company's commercial API in violation of its terms to mass-harvest responses, then feeding those responses into training pipelines for a rival system.
The concern is economic as much as technical. Frontier labs spend tens to hundreds of millions of dollars on data, compute and alignment to build models like Claude Opus 4.1 and Claude Sonnet 4.5. If a rival can replicate much of that capability for the cost of API calls, it undercuts the incentive to innovate and creates major security risks.
How The Alleged Attacks Worked
According to Anthropic, the three campaigns shared a similar playbook but operated separately:
Mass account creation: Operators allegedly spun up thousands of new Claude accounts, often using disposable emails, automated signup tools and VPNs and cloud infrastructure to mask locations and dodge rate limits and fraud detection.
Evasive prompting: Instead of asking straightforward questions, the accounts allegedly used jailbreaks, roleplay scenarios, and prompts specifically designed to force Claude to reveal its extended chain-of-thought reasoning. Anthropic said attackers were particularly interested in coding, math, agentic tool use and refusal boundaries.
Targeted harvesting: The report describes highly structured querying — tens of thousands of near-identical prompts with small variations — consistent with building large synthetic datasets for supervised fine-tuning and reinforcement learning, rather than normal user behavior.
Anthropic said it linked the activity to the three labs through a combination of technical indicators, operational patterns and supporting intelligence, and that it has since banned the associated accounts and hardened its API abuse detection.
Why Anthropic Went Public Now
The timing is no accident. The past nine months have seen a brutal acceleration in the AI race.
DeepSeek stunned Western labs in early 2025 with its highly efficient R1 reasoning model, followed by rapid releases from Alibaba's Qwen team and Moonshot AI's Kimi models that closed the gap with U.S. systems on coding and agent benchmarks. At the same time, Washington has tightened export controls on advanced AI chips to China, pushing Chinese labs to seek efficiency gains through software and training-data advantages.
Anthropic has previously warned about distillation attempts without naming names. In January, OpenAI and Microsoft said they had investigated DeepSeek-linked accounts for possible improper distillation of OpenAI models. By naming Alibaba, Moonshot AI and DeepSeek directly, Anthropic is breaking with that cautious approach.
Analysts say the public attribution serves three goals: deterring future abuse, pressuring cloud and API resellers to enforce controls, and making the case to policymakers that model weights and outputs need stronger protection as strategic assets.
What The Accused Labs Stand To Gain
For Chinese labs operating under chip constraints, distillation offers a shortcut.
Copying advanced reasoning traces can dramatically improve a smaller model's performance on complex tasks without requiring massive new human-labeled datasets. It can also help map a competitor's safety filters — learning exactly where Claude refuses and where it complies — and then train to replicate or bypass those behaviors.
Anthropic alleges the harvested data was likely used to improve the labs' own flagship models, including Qwen, Kimi and DeepSeek series, all of which have made notable leaps in reasoning and agentic capabilities this year.
Security experts note that proving a shipped model was trained on distilled data is extremely difficult after the fact, which makes API-level disruption the primary line of defense.
What This Means For AI Security
The report signals a shift in how frontier labs think about security. The threat is no longer just stolen weights or leaked system prompts, but slow, distributed siphoning via legitimate-looking API traffic.
Expect Anthropic and rivals to respond with stricter know-your-customer checks for API access, tighter limits on reasoning-token visibility, more aggressive watermarking and dataset tracing, and faster account bans for anomalous usage patterns.
For enterprise users, that could mean more friction: additional verification, new rate limits on bulk inference, and less access to raw chain-of-thought outputs.
For the industry, it raises an uncomfortable question: in an era where any public API can become a training source for competitors, how open can frontier models afford to be?
The Bigger Picture: AI Rivalry Turns Confrontational
Beyond the technical details, Thursday's report marks a new phase in U.S.-China AI rivalry — one fought not just over chips and talent, but over data provenance and intellectual property.
If Anthropic's allegations hold, they could fuel calls in Washington for stronger enforcement against terms-of-service violations as a national security issue, and complicate partnerships involving cloud providers, resellers and open-model distribution.
They could also trigger retaliation in narrative if not in policy, with Chinese firms portraying the accusations as an attempt to stifle competition as their models gain global traction.
What to watch next is whether OpenAI, Google DeepMind and Meta corroborate Anthropic's findings, whether the named Chinese labs publish technical rebuttals or legal challenges, and whether this leads to industry-wide standards for detecting and disclosing distillation abuse.
One thing is clear: the distillation wars are no longer a behind-the-scenes worry. They are now front and center in the battle for AI supremacy.
Get All The Latest Updates Delivered Straight To Your Inbox For Free!