Anthropic Accuses Alibaba of Copying Claude
Anthropic says Alibaba copied Claude across 28.8 million questions. But can you really steal a model just by asking it things?

28.8 million questions put to Claude in six weeks. No server breach, no source-code theft, no leak. Just conversations, the kind millions of people have with an AI assistant every day. And yet this is what Anthropic calls the largest attack ever waged against it.
In a letter sent to the White House and to US senators on June 10, Anthropic accuses Alibaba of using nearly 25,000 fake accounts to query Claude between April 22 and June 5, 2026. The presumed goal: siphon off the model's know-how to train a competitor at a fraction of the cost. The term used is "adversarial distillation." The word that runs through the press coverage is "theft."
What Anthropic alleges, and against whom
According to Anthropic, the operators were tied to Alibaba and to Qwen, its AI lab. The targeted capabilities were anything but trivial: agentic reasoning, code generation, long-horizon tasks that require the model to hold a thread across several steps. In other words, what makes Claude genuinely valuable today, not the easy answers.
The script is familiar. In February 2026, OpenAI made the same case before Congress about DeepSeek, accused of distilling its models through masked third-party routers. As early as January 2025, OpenAI and Microsoft already suspected that DeepSeek R1 had been partly trained on ChatGPT outputs. Anthropic itself had named DeepSeek, Moonshot, and MiniMax a few months earlier. Every time, a US lab accuses a Chinese lab, and asks Washington to act.
Distillation, explained without the jargon
Picture a sketch artist who can never see the suspect. They interview hundreds of witnesses, cross-check the descriptions, and end up drawing a face close enough to be recognized on the street. They never met the person. They just asked a lot of questions to the people who did.
That is how distillation works. You touch neither Claude's code nor its weights, the billions of parameters that actually make up the model and that stay locked inside Anthropic. You only observe its answers, in bulk, and train another model, the student, to reproduce them. The teacher never even knows class is in session.
And this is where it gets awkward for Anthropic. Distillation is not some clandestine technique. It is a routine method, used all the time to build smaller and cheaper models. It only becomes contentious in one specific case: when it serves to replicate a frontier model, and when it runs through fraud. Here, the 25,000 fake accounts.
The real issue: you don't own what you think you own
Put the legal question bluntly. If Alibaba did what it is accused of, what exactly did it steal, in the eyes of the law?
Not the code, there was no break-in. Not the model's weights, they never left Anthropic. What remains are the outputs, the answers Claude generated. And in the US, the Copyright Office holds that text produced without a genuine human author is not eligible for copyright protection. Several legal analyses draw an uncomfortable conclusion for the labs: if API outputs are not covered by copyright, reusing them to train a model, even at scale, probably does not amount to infringement.
So what Anthropic is left with is not ownership of the model. It is the contract. The terms of service forbid using Claude to build a rival model, and creating 25,000 fake accounts to bypass the safeguards breaches that contract. The fight is not playing out on intellectual-property turf, but on fraud and unfair competition. The distinction matters: this is not about guarding a treasure, it is about enforcing access conditions.
The accuser has a short memory
There is an irony hard to ignore. The models now crying foul over the theft of their outputs were themselves trained on vast quantities of text, images, and code scraped from the web without asking anyone's permission. Declic documented it in April: the same industry that helped itself without consent to build itself now fiercely defends the fruit of that build. See our piece OpenAI, Anthropic, Google and the hypocrisy of scraping.
The line between tolerated distillation and "illicit" distillation does not rest on the technique, which is identical in both cases. It rests on scale and intent. A few thousand requests to cobble together a small model, and nobody blinks. Twenty-eight million to reconstruct a direct competitor, and the vocabulary shifts to theft. It is a dial, not a clean line, and each side slides it depending on whether it is accusing or defending.
What you actually protect
Alibaba did not respond to press inquiries at the time. The investigation is still to be run, and the accusation, however well-quantified, remains an accusation.
But the case says something beyond this one episode. When you spend billions to train a frontier model, you are not buying a safe. You are buying a head start, a barrier to entry that no one else can afford to clear.
The trouble is that this barrier wears down with every request. The more useful the model, the more it answers, and the more it teaches anyone who knows how to listen. The value of a frontier model rests on the idea that it is too expensive to reproduce. This story suggests that sometimes, all it takes is asking enough questions.
Topics covered:
Frequently asked questions
What is model distillation in AI?
What exactly is Anthropic accusing Alibaba of?
Is distillation illegal?
Can Anthropic sue for intellectual property theft?

Alexandre Noto
Co-founder & Tech Expert
Alexandre has been in tech for over 20 years. Entrepreneur, software architect and AI enthusiast, he translates complex concepts into accessible explanations. At Declic Media, he is the technical voice that makes AI understandable for everyone.
All articles by Alexandre →