Anthropic released a report detailing a surge in distillation attacks against its models, specifically Claude. The report indicates that unauthorized labs have intensified efforts to extract the chain of thought from Claude’s responses, a technique used to train smaller models. Nearly 200 million exchanges linked to these attacks were observed across five campaigns.
The largest campaign originated from Alibaba, involving 151 million exchanges between May and July 2026. These exchanges, utilizing a single fixed prompt, were attributed to the company’s Qwen family of models. The campaign peaked at nearly three million exchanges per day.
Another significant campaign stemmed from Moonshot AI, the manufacturer of Kimi. This campaign routed requests directly to Claude, primarily targeting the company’s Opus model. Over a 10-day period, nearly 300,000 requests were made through a network of 5,000 accounts, with one specific request involving the assessment of closed-circuit surveillance footage.
Anthropic describes these attacks as aggressive and sophisticated, highlighting the potential for misuse of extracted chain-of-thought data. The company’s typical approach of displaying summarized thinking blocks was circumvented, allowing attackers to directly access the model’s reasoning processes. This underscores the importance of robust defenses against such attacks, particularly as competition in the AI space intensifies.



