Society & safeguards

Anthropic, AI's Self-Proclaimed Security Champion, Just Leaked Its Own Secrets. Twice in 5 Days.

4 min read

The company that positions itself as AI's most responsible player just exposed a model it calls an 'unprecedented cybersecurity risk,' then leaked the source code of its flagship product. Five days apart.

The free AI newsletter
Anthropic, AI's Self-Proclaimed Security Champion, Just Leaked Its Own Secrets. Twice in 5 Days.

Anthropic exists because its founders thought OpenAI wasn't taking security seriously enough.

Dario and Daniela Amodei left OpenAI in 2021 to build a company with a two-word pitch: safety first. Every announcement, every blog post, every interview hammers the same message: we're building the world's most powerful AI, but we're doing it responsibly.

Last week, over the span of five days, Anthropic leaked its own secrets. Twice.

Leak #1: The model that 'poses unprecedented cybersecurity risks'

On March 26, Fortune discovered roughly 3,000 internal Anthropic files sitting in an unsecured public database.

Among those files: a draft blog post announcing "Claude Mythos," internal codename Capybara. It's Anthropic's next flagship model, described as a "step change" and the most powerful model the company has ever built. Stronger than Opus 4.6, their current best.

The kicker: in this draft, Anthropic writes in plain text that Capybara "poses unprecedented cybersecurity risks" and "ushers in a wave of models capable of exploiting vulnerabilities at a speed that outpaces defenders."

In summary: Anthropic is preparing a model it judges dangerous, and the entire world learns about it because someone forgot to put a password on the database.

The next day, cybersecurity company stocks dropped.

Leak #2: The source code of Claude Code

Five days later, on March 31, a security researcher discovered the complete source code of Claude Code available in plaintext on npm, the package registry developers use to distribute software.

500,000 lines of code. 1,900 files. The entire system that powers Claude Code: the agent architecture, internal APIs, guardrails, system instructions, unreleased features.

The cause: a 60-megabyte sourcemap file accidentally included in the published package. Someone didn't remove a debug file before shipping. This is the kind of mistake interns make.

The code was extracted and published on GitHub within hours. It will never disappear.

Worse: this is the second time in a year Claude Code has leaked the same way. February 2025, identical scenario.

Anthropic's response: "This is not a security vulnerability, it's a packaging issue caused by human error." As if human error isn't precisely the problem.

What the leaks are not

Let's be honest about what did NOT leak.

The model weights, the AI's "brain," are intact. No customer data was exposed. No credentials were compromised. The first leak is mostly embarrassing, the second is more serious but not catastrophic.

Other companies have similar leaks. Epic Games, Nintendo, Google, Tesla have all left files lying around in public places.

But none of those companies make security their primary selling point.

The structural irony

This is where it gets genuinely interesting.

The leaked blog post wasn't a random internal memo. It was a text prepared to show how seriously Anthropic takes security. The message was: "Our new model is so powerful we're testing it with cybersecurity experts before releasing it."

And that text leaked because it was sitting in a public database without a password.

Roy Paz, security researcher at LayerX: "Normally, large companies have strict processes and multiple checks before code reaches production, like a vault that requires several keys to open. At Anthropic, it seems this process wasn't in place."

Anthropic generates $19 billion in annualized revenue. They can't manage to remove a debug file from an npm package.

What this says about the industry

The problem extends beyond Anthropic. If the most vocal company about AI safety makes these kinds of errors, twice in five days and twice in a year for the same product, what can we reasonably expect from everyone else?

The AI industry is moving faster than its capacity to secure itself. Models become more powerful, agents more autonomous, the stakes higher. And the basic processes, the ones that prevent a file from ending up on the internet, aren't keeping pace.

This isn't a technology problem. It's a hygiene problem. And hygiene can't be bought with $19 billion in revenue.

Next time an AI company promises your data is secure, remember that Anthropic couldn't even protect its own.


Sources:

Topics covered:

SecurityAnalysis

Frequently asked questions

What did Anthropic leak?
Two leaks in 5 days: an internal blog post revealing their next model Claude Mythos (described as an 'unprecedented cybersecurity risk'), then the complete source code of Claude Code (500,000 lines).
Did the Claude Mythos model itself leak?
No. The model weights (the AI's 'brain') did not leak. Only an internal blog post describing its capabilities and a CEO event were exposed.
Was customer data compromised?
No. Anthropic confirms no customer data or credentials were exposed. The leaks involved internal code and strategic documents.
Alexandre Noto

Alexandre Noto

Co-founder & Tech Expert

Alexandre has been in tech for over 20 years. Entrepreneur, software architect and AI enthusiast, he translates complex concepts into accessible explanations. At Declic Media, he is the technical voice that makes AI understandable for everyone.

All articles by Alexandre →
The free AI newsletter