Power & systems

OpenAI, Google, Anthropic: The Scraping Hypocrites

5 min read

OpenAI, Anthropic and Google have formed an anti-scraping coalition, despite having built their models by scraping the entire web. The irony writes itself.

The free AI newsletter
OpenAI, Google, Anthropic: The Scraping Hypocrites

On April 6, 2026, three companies locked in total commercial warfare did something unusual: they formed an intelligence coalition. OpenAI, Anthropic and Google are now sharing information via the Frontier Model Forum to detect what they call "adversarial distillation" -- competitors systematically extracting their model outputs to train their own without paying.

What are they trying to protect? Exactly what they built by pillaging the rest of the web.

What is actually happening

The technique is straightforward. You create 24,000 fake accounts (the figure documented by Anthropic alone). You query GPT-4 or Claude millions of times. You record every exchange. Then you use that dataset to train your own model, absorbing years of RLHF, alignment, and fine-tuning work, without paying a cent or obtaining any authorization.

Anthropic says it documented 16 million such unauthorized exchanges from three Chinese companies: DeepSeek, Moonshot AI, and MiniMax. The scale is serious enough that three commercial rivals have temporarily paused their war with each other.

The sharing mechanism? The Frontier Model Forum, a body created in 2023 by OpenAI, Anthropic, Google and Microsoft to "promote AI safety." In practice, it is a self-regulated cartel that sets its own rules with no external legal mandate. We'll come back to that.

The contradiction nobody names

There is an interesting detail in this story.

These same companies are currently being sued by newspapers (New York Times, Chicago Tribune), authors, musicians, and publishers, for scraping their work without permission to train their models. Zero compensation. Zero authorization. Just: "it's research, it's fair use."

The New York Times documented thousands of its articles reproduced almost word for word in GPT-4 outputs. That lawsuit is still ongoing.

So here is the official position of the Big Three in April 2026:

  • Scraping human authors' data to build multi-billion dollar models: fair use, innovation, freedom of research.
  • Scraping their models' outputs to build competing models: theft, ToS violation, a hostile act requiring coordinated international response.

The distinction is legally defensible: ToS agreements are contracts, training data copyright is a separate question. But morally, the contortionism is spectacular.

The FMF: the referee who plays on the team

What makes this story even more interesting is the structure chosen for this "coalition."

The Frontier Model Forum is not a regulator. It is not an independent institution. It is a body created and controlled by the four largest AI companies in the world, which has appointed itself as arbiter of access to AI. No democratic mandate. No external oversight. Rules set by the direct beneficiaries.

The FMF will now share "intelligence" on actors attempting to distill their models -- meaning: decide who gets API access and who does not. A coordinated ban by three players who control a massive share of the global AI infrastructure.

For European startups or academic researchers using the APIs of these three labs, this is worth noting: access rules are now set by a private cartel, based on geopolitical criteria as much as contractual ones.

The second-order effect everyone ignores

The immediate reaction is understandable. If DeepSeek is vacuuming Claude's responses to train its own model, that is a real problem, not just ethical but competitive. A model trained on aligned model outputs without the alignment work is potentially dangerous.

But the announced countermeasures go beyond detecting bad actors. Tighter APIs also mean:

  • Higher access costs for independent developers
  • Stricter rate limits for academic research
  • A higher barrier to entry for non-US players

The practical result: increased power concentration around players who have the resources to train models without depending on the Big Three's APIs.

The irony is complete. In response to adversarial distillation, the major labs will likely tighten their APIs, which benefits exactly the players who can do without them (Mistral, serious open-source models, Meta). And hurts the smaller players who were building legitimate products on those APIs.

What this tells us about regulation

There is a gaping legal void here.

Large-scale adversarial distillation is, as of today, not illegal in most jurisdictions. It is a violation of Terms of Service: a civil contract, not a criminal offense. The only current lever is API exclusion and account banning.

But if these techniques can genuinely bypass years of safety and alignment work to create dangerous models, then this is a national security issue, not just an IP question. And a national security issue cannot be managed by a cartel of private companies regulating themselves.

The EU has the AI Act. The US has non-binding memos and the FMF. Meanwhile, the race continues.

What to take away

OpenAI, Anthropic and Google are protecting themselves against exactly what they did to get where they are. That is not a moral scoop: companies defend their interests, that is their job.

What is interesting is the structure they chose to do it: a self-regulated cartel, without legal mandate, making decisions with global consequences for access to AI.

The FMF was not built to guarantee fair and safe access to AI. It is judge and jury in its own case.


Sources: Bloomberg (Apr. 6, 2026), The Japan Times (Apr. 7, 2026), Built In, MIT Technology Review, Euronews

Topics covered:

RegulationOpenAIAnthropic

Frequently asked questions

What is adversarial distillation?
Adversarial distillation is the process of querying an AI model at massive scale using fake accounts, recording its responses, and using that data to train a competing model. Anthropic documented 16 million unauthorized exchanges of this type.
Why are OpenAI, Google and Anthropic forming an anti-scraping coalition?
The three labs are now sharing intelligence via the Frontier Model Forum to detect and block actors who systematically extract their model outputs to train competing systems, particularly Chinese companies.
Why is this coalition contradictory?
These same companies are being sued by publishers and authors for scraping their work without permission to build their own models. They are now demanding the protections they never extended to creators.
What is the Frontier Model Forum?
The Frontier Model Forum is a body created in 2023 by OpenAI, Anthropic, Google and Microsoft. It is not an independent regulator: it is a self-regulated cartel, with no external legal mandate, controlled by its founding members.
What does this mean for independent developers?
Tighter APIs mean higher access costs, stricter rate limits for academic research, and a higher barrier to entry for non-US players.
Alexandre Noto

Alexandre Noto

Co-founder & Tech Expert

Alexandre has been in tech for over 20 years. Entrepreneur, software architect and AI enthusiast, he translates complex concepts into accessible explanations. At Declic Media, he is the technical voice that makes AI understandable for everyone.

All articles by Alexandre →
The free AI newsletter