Society & safeguards

How AI Screwups Get Counted: By Reading Angry Posts on X

6 min read

Are there more incidents, or just more complaining about them?

The free AI newsletter
How AI Screwups Get Counted: By Reading Angry Posts on X

Over 300 AI Screwups in July, and One Big Question

Over 300 loss-of-control incidents in July, nearly double June's count, and more than 1,600 since the start of the year. The number comes from the Loss of Control Observatory, a tracker run by the Centre for Long-Term Resilience and funded by the UK AI Security Institute's Challenge Fund. It was shared with the Guardian, which reported it on August 29.

It's the only public effort trying to measure what AI agents actually do once they're turned loose on real users. Its raw material fits in a single sentence: the screenshots and conversation links people post on X when their agent has just done something weird.

Which raises the obvious question. To land on 300 incidents in July, how many posts had to get sifted through? That answer doesn't exist anywhere.

Why This Tracker Exists

The observatory was announced on February 2, 2026, and the page announcing it puts the word prototype right in the title. It starts from a complaint few people would dispute: existing incident databases run on voluntary reports or press coverage, arrive days late, and never go looking for concealment behavior.

The criticism also lands on the labs. Speaking to the Guardian, Tommy Shaffer Shane, a senior policy manager at the CLTR, argues that companies don't necessarily monitor their own internally deployed models, and wants them to report even minor incidents.

The instinct is sound, and the gap it's aimed at is real. It's the method that deserves a second look, because it amounts to measuring car crashes by counting the drivers who post a photo of their bumper.

A Two-Stage Filter, One Stage Run by an AI Judge

The CLTR published its methodology on March 27, in a 76-page report that remains its last word on AI. It lays out the full funnel, covering October 12, 2025 through March 12, 2026.

3,391,950 posts collected on X, kept because they mention an AI, a problem, and include an image or a conversation link. A first language model, fast and cheap, keeps 183,420 of them: its instructions explicitly tell it to cast a wide net and lean toward inclusion whenever in doubt. A second model scores what's left. After deduplication: 698 incidents.

That second model is Claude Opus 4.6. The CLTR tested nine models and measured how well each agreed with two human reviewers: the two humans agreed with each other at 0.70; Opus 4.6 agrees with them at 0.77. So the judge deciding what counts as an AI going rogue is itself an AI, and it agrees with the human reviewers more than the human reviewers agree with each other.

What Actually Counts as an Incident

The counting bar sits at 5 out of 9. The report spells out exactly what that 5 is worth: the evidence can be only partial, the behavior can be relatively minor, and there can be doubts about how credible the witness account even is. That's the floor of the category, not its ceiling.

What clears the sieve is sometimes genuinely serious. The March report describes Google's Antigravity coding environment which, asked to clear a cache, ran a delete command on the root of a user's D: drive and wiped out years of photos and client work without asking for confirmation. It also describes Claude Code running a terraform destroy that took down a production infrastructure holding two and a half years of student coursework, and an agent managing a token treasury that sent roughly $270,000 to a stranger, followed by a 60% price crash.

The CLTR adds the caveat that actually matters: the damage is real, but nothing proves it reflects deliberate strategy rather than a badly understood instruction. The line between an agent that follows orders poorly and one chasing a different goal entirely, the report notes, is inherently blurry.

The Missing Number

In March, the CLTR does the thing that makes its figure legible. It doesn't just announce a 4.9x jump in incidents between the first and last month of the window. It stacks that against growth in chatter on the same topic, which was only 1.7x, and against growth in negative AI discussion generally, 1.3x. Incidents are climbing noticeably faster than the background noise, and it's that comparison, and only that comparison, that earns the word signal.

The report goes further and spells out the competing explanation in plain terms. This rise, it says, could just as easily come from wider agent deployment, heavier usage, or a shift in reporting habits.

None of that means nothing is happening. If people are talking about it more, there's probably more of it happening. But a number that grows because the fleet of agents grows doesn't mean agents are misbehaving more often, only that there are more of them around to misbehave. A volume isn't a rate, and it was the March comparison that let anyone tell the two apart.

For July, none of those guardrails ship with the number. Not the count of posts collected, not the count kept after pre-filtering, not the growth in chatter. A doubling in reports says nothing on its own if you don't know whether the population doing the reporting doubled too.

A Number That Was Shared, Not Published

The Guardian is upfront about this: the latest results were shared with it directly. They appear in no CLTR publication. The organization's AI page lists four publications, the most recent dated March 27, and nothing since then touches the observatory.

Lining up the two documents raises a question nobody can settle. The last published month, February 9 to March 12, counted 319 incidents. July is announced as "more than 300," phrasing that hands you a floor without handing you the number. The Guardian's headline calls it a record, but the public data can't confirm that or rule it out.

One more detail says more than any press release could. On July 21, right in the middle of the month whose count is said to have doubled, the CLTR posted a job listing for an engineer to rebuild the observatory's collection pipeline and classification system. Among the expected deliverables: exportable datasets and, eventually, an API letting outside parties pull the data themselves. That listing described the observatory as a pilot.

Two Counters in Three Days

Two days ago, we took apart the 700-agent figure that made the rounds after the Hugging Face incident: it wasn't counting 700 durable entities but agent runs over a few days, and it came out of an automated classifier. The mechanism is different here, but the outcome rhymes.

None of this says agents are behaving themselves: the March report shows the opposite, transcripts included. What it does say is that the only instrument we have for counting them also measures, and maybe mostly measures, how many people felt like posting about their bad day in public. The observatory said as much itself, in March. The July number traveled without that line attached.

Topics covered:

SecurityAnalysis

Frequently asked questions

What is the Loss of Control Observatory?
An AI agent incident tracker run by the Centre for Long-Term Resilience and funded by the UK AI Security Institute's Challenge Fund. The page announcing it, dated February 2, 2026, carries the word prototype in its title. Its raw material: the screenshots and conversation links that users post on X.
Where does the 'more than 300 incidents in July' figure come from?
It was shared with the Guardian, which reported it on August 29. It doesn't appear in any CLTR publication: the CLTR's last publication on AI dates from March 27, 2026. So no outside party can verify it.
How is an incident counted?
Between October 12, 2025 and March 12, 2026, 3,391,950 posts were collected on X, of which 183,420 were kept by a first language model, then scored by a second one, Claude Opus 4.6. After deduplication, 698 incidents remain. The counting threshold is 5 out of 9, a floor that allows for only partial evidence, a relatively minor behavior, or doubts about how credible the testimony is.
Are AI agents actually going rogue more often?
For the published period, the observatory itself published the comparison that justifies talk of a signal: incidents rose 4.9x between October 2025 and March 2026, while the volume of discussion on the topic only grew 1.7x. The report also names the competing explanations: wider agent deployment, heavier use, changing reporting habits. For July, that comparison was never published.
Does the July figure beat the March one?
Nothing public says so. The last published month, February 9 to March 12, counted 319 incidents; July is announced as more than 300, a lower bound that doesn't give the actual number. The Guardian's headline calls it a record, but the public data can neither confirm nor rule that out.
Katja Liersch

Katja Liersch

Co-founder & Journalist

Katja is a journalist and TV producer. With decades of experience in mainstream media, she brings to Declic Media the perspective of those discovering AI: curious, demanding and pragmatic. She ensures every piece of content truly speaks to everyone.

All articles by Katja →
The free AI newsletter