Section

Society & safeguards

Security, privacy, ethics: what AI does to society, and the safeguards we give it — or don't.

Security · Privacy · Ethics51 articles

Society & safeguards

No, the AIs Didn't Escape. The Locks Were Already Off

Three labs disclosed agent incidents in ten days. Two of them trace back to the same testing vendor. The third starts with a task nobody could solve.

Alexandre Noto · Aug 6, 2026 · 6 min

The Verified Number

AI Writes Half Your Code? Nobody Actually Measured That

This week's most-shared AI number didn't come from a code repository. It came from a survey asking developers what they believe.

Alexandre Noto · Aug 3, 2026 · 6 min

Society & safeguards

Anthropic's Test Environment Wasn't Isolated. Real Firms Got Hit

Three Anthropic models reached real company systems during tests that were supposed to be cut off from the internet. The flaw was in the test setup itself.

Alexandre Noto · Jul 31, 2026 · 6 min

Society & safeguards

Hugging Face: Guardrails Blocked the Defenders, Not the Attacker

Hugging Face's post-mortem walks through three AI safety guardrails. None slowed the attacker down. All three got in the defenders' way.

Alexandre Noto · Jul 30, 2026 · 5 min

Society & safeguards

The Hugging Face breach came from OpenAI

The agents that broke into Hugging Face weren't hackers. They were OpenAI's own models, mid-evaluation, cheating on a benchmark.

Alexandre Noto · Jul 22, 2026 · 5 min

Society & safeguards

An OpenAI Model Found a Way Around Its Own Sandbox

OpenAI disclosed that one of its unreleased models broke out of its sandbox during a test. What's reassuring about the story is exactly what's unsettling.

Alexandre Noto · Jul 21, 2026 · 5 min

Society & safeguards

Did OpenAI Know GPT-5.6 Deletes Files?

OpenAI documented that GPT-5.6 deletes files without consent, then shipped it thirteen days later. When transparency becomes a shield.

Alexandre Noto · Jul 18, 2026 · 5 min

Society & safeguards

Kids and AI: What Convenience Takes Away

A study on Google's AI and minors reveals one mechanism behind two problems we tend to file as unrelated: the removal of friction.

Alexandre Noto · Jul 15, 2026 · 5 min

Society & safeguards

J-Space: Where Claude Thinks Without Saying It

Anthropic says it can read an internal space where Claude holds thoughts it never writes down. Here is what it changes for AI transparency.

Alexandre Noto · Jul 7, 2026 · 5 min

Society & safeguards

The AI That Screens Your CV on Its Own Is Illegal in Europe

The robot that rejects your CV alone, with no recourse, is largely a fantasy and mostly illegal in Europe. The sense of injustice, though, is very real.

Alexandre Noto · Jul 2, 2026 · 5 min

Society & safeguards

When the human watching the AI checks out

Amazon admits the human-in-the-loop, the official safety net for AI, rests on an attention nobody actually sustains.

Alexandre Noto · Jun 22, 2026 · 5 min

Society & safeguards

Do Chatbots Erode Critical Thinking?

A study links heavy chatbot use to weaker critical thinking, especially among young people. The correlation is real. The causation, far less so.

Alexandre Noto · Jun 19, 2026 · 5 min

Society & safeguards

37 dark patterns in AI chatbots: what the CDT report reveals

The Center for Democracy and Technology published on May 28 a taxonomy of dark patterns in ChatGPT, Claude, Gemini, Replika, and Character.AI.

Alexandre Noto · Jun 9, 2026 · 6 min

Society & safeguards

When the fake says what the real ones no longer dare to

An AI-generated speech attributed to Namibia's first female president has circulated since October 2025. The presidency denied it seven months ago. People keep sharing it.

Alexandre Noto · Jun 6, 2026 · 6 min

Society & safeguards

Doctolib, AI, and the silence of France's major newspapers

Five specialist outlets covered Doctolib's April 2026 privacy policy. None of France's major dailies picked up the story 72 hours later.

Alexandre Noto · Jun 5, 2026 · 6 min

Society & safeguards

Anthropic Expands Mythos to 200 Orgs as It Warns of 100M Risk

Anthropic expanded Mythos access to 150 new partners and wrote in the same post that a breach would affect 100 million people.

Alexandre Noto · Jun 4, 2026 · 6 min

Society & safeguards

Someone Took Over Obama's Instagram by Politely Asking Meta's AI Bot

Last weekend, Meta's AI support bot handed OTP codes to attackers who simply asked for them. The Obama-era White House account, the Space Force, and Sephora all got defaced. A human at the support desk would have said no.

Alexandre Noto · Jun 2, 2026 · 6 min

Society & safeguards

They left AI agents to run a city for fifteen days: three out of four collapsed

Emergence AI handed five virtual cities to AI agents for fifteen days. Grok was dead in four, GPT in seven, Gemini torched city hall. The marketing pitch for agent societies does not survive contact with the experiment.

Alexandre Noto · Jun 1, 2026 · 5 min

Society & safeguards

Meta spied on its employees to train the AI replacing them

Leaked April 30 audio: Zuckerberg explains the Model Capability Initiative to employees. Three weeks later, 8,000 of them are laid off.

Alexandre Noto · May 26, 2026 · 6 min

Society & safeguards

Grok deepfakes and SpaceX IPO: what French prosecutors are really telling the SEC

Paris prosecutors suspect the Grok scandal was orchestrated to inflate X and xAI valuations ahead of the SpaceX IPO on June 12.

Alexandre Noto · May 21, 2026 · 5 min

Society & safeguards

Palantir lands at NHS and ICE on the same day

Two continents, the same day, the same company. Here is what Palantir actually does and why 2026 changes the regulatory equation.

Alexandre Noto · May 12, 2026 · 5 min

Society & safeguards

Chrome installed 4 GB of AI on your machine and quietly removed the line that promised your data would stay put

Five words vanished from Chrome 148's settings, right when Gemini Nano was being silently pushed to a billion devices.

Alexandre Noto · May 11, 2026 · 5 min

Society & safeguards

When an AI Overview cancels your gig

On May 4, 2026, a Canadian fiddler is suing Google for CA$1.5 million. AI hallucinations are no longer a private problem.

Alexandre Noto · May 5, 2026 · 5 min

Society & safeguards

Claude wiped a production database in 9 seconds. Then wrote an apology.

Claude Opus 4.6 vaporized PocketOS's production database via Cursor. One GraphQL mutation, 9 seconds, three months of data gone.

Alexandre Noto · May 1, 2026 · 5 min