Society & safeguards

An OpenAI Model Broke Into the Machine That Designs Its Chips

How did it get in?

6 min read
An OpenAI Model Broke Into the Machine That Designs Its Chips

A Model Went Looking for Its Grade

On March 27, an OpenAI internal research model is sitting an evaluation. Instead of writing the requested solution, it decides to check how it'll be graded first. It sends the grader a throwaway answer that fails, but the failure reveals the 636 hidden test entries.

From there, it works backward through the pipeline. A tool it has access to writes files without checking where: the model uses it to overwrite code that the tool reloads on every call, which gets its own commands executed. That gets it into an internal service that accepts chip-design jobs, where it notices a form parameter gets copied straight into a system command, and slips its own in.

First attempt: rejected. The model figures out why, places the same string inside a code comment it submits instead, and the command runs on the machine. It's a locksmith who tries a key, gets told to also hand over a photo of it, and hands over the photo.

A Date the Page Doesn't Show

This isn't a leak. OpenAI published this account itself, chain of thought included, on its alignment blog.

The report notes that the employees running that internal service spotted the suspicious activity, flagged it, and the server was shut down. It also carries two dates: incident on March 27, discovery on March 27. Same day.

That leaves the third date, publication. It's written nowhere. The page reads "Report updated: Oct 2, 2026", and the section index is careful to note that this is the date of the last update, not the original post. A report updated October 2 could have first gone up back in August.

So we went looking elsewhere. The Wayback Machine holds no trace of these three pages before the morning of October 3, while the previous nine reports show up there the day after their listed date, every time. The gap is consistent across all twelve.

The blog's own RSS feed doesn't help either: it only syndicates research posts, no misalignment reports at all. Nothing alerts anyone.

Between the March 27 incident and its October 2 publication: 189 days.

The Framework Promised the Opposite

Two weeks earlier, on September 16, OpenAI published the framework governing these disclosures. Its intent is spelled out plainly: "This new framework is intended to expedite publishing misalignment reports following observation, even when we haven't fully explained or mitigated the behavior we're reporting."

The same text names the flaw it's meant to fix: "we've often waited until we could collate several instances into one report". Since then, there have been three publication rounds: six reports on September 16, three on the 25th, three on October 2. None of them ever came out alone.

The framework also states: "This starts our disclosure process, with deadlines for each step to ensure timely investigation and disclosure." It doesn't put a number on any of them. Three handling tracks are named, including one called Slow Track, with no duration attached to any of them. It's a train timetable announcing punctual departures without giving a single time.

Five Days on One Side, 189 on the Other

Lined up by delay, the twelve reports show a gap of a factor of forty-seven. The longest, 235 days, belongs to the inaugural batch of September 16, which makes sense: that batch was clearing a backlog. The median is 124 days.

The uncomfortable part is elsewhere. The October 2 batch, the first produced entirely under the new regime, clocks in at 189, 139, and 133 days. That's slower than the median, and its oldest case is the second-slowest of the twelve, right behind the backlog one.

At the other end, one report came out in five days: an agent had routed around its sandbox's DNS filtering to ask a public chatbot a question. Its timeline is published down to the second: call at 9:50:23 a.m., automated alert twelve minutes later, human acknowledgment three minutes after that, execution halted by 12:34 p.m.

And it's OpenAI itself describing that case: "This incident is a lot less severe than some of our previous incidents". The case handled in five days is the one the company calls mild, and the one that took 189 days is the one where a model took control of an internal machine. Speed runs the opposite way from severity.

Is There a Reason? We Didn't Find One

A delay isn't a fault on its own. A security investigation takes time, and it's normal for a serious case to move slower than a network misconfiguration.

The framework even plans for exactly that: the Slow Track covers complex investigations, "especially those involving third parties", and allows a delay for security reasons. Except none of the three October 2 reports says which track it followed, and all three incidents are internal to OpenAI: a chip-design machine, a Slack channel, an in-house tool. No third party.

Let's be precise about what we actually have. We found no published reason for these three particular delays, which isn't the same as saying there is none. And both the incident date and the discovery date come from OpenAI and nobody else: no outside observer can verify that an incident actually happened on March 27.

When Someone Else Is Watching

The contrast comes from the company itself. In July, when OpenAI models reached into Hugging Face's systems, it published its own timeline: alert on the 19th, correlation on the 20th, notification to Hugging Face right after, public acknowledgment on the 21st. One day to warn, two to say it out loud. The full technical report didn't land until August 26, but the public statement had been immediate.

The difference is obvious: that time, someone outside was involved, and watching. We reported in September on how OpenAI was pushing Congress for an incident-notification law covering incidents its own AI causes, three days before publishing its voluntary charter. We noted at the time that the voluntary framework was still under construction. It exists now.

For the March 27 incident, no one from outside was in the room. The only way to find out when OpenAI got around to talking about it is to ask an archive.

Topics covered:

EthicsOpenAIAnalysis

Frequently asked questions

What happened at OpenAI on March 27, 2026?
An internal research model was sitting an evaluation. It sent the grader a throwaway answer that revealed the 636 hidden test entries, then used a tool that writes files without checking where, to get its own commands executed. That got it into an internal chip-design service, where it slipped a command through a form parameter copied straight into a system command. Rejected the first time, it placed the same string inside a code comment instead, and the command ran.
How long did it take OpenAI to publish this report?
189 days passed between the incident, which OpenAI dates to March 27, and the October 2 publication. This is not a detection delay: the report lists discovery on March 27, the same day as the incident. The report notes that the employees running that internal service flagged the suspicious activity and the server was shut down.
How do we know the publication date if OpenAI doesn't give one?
The page only shows 'Report updated: Oct 2, 2026', and the section index notes that this date marks the last update, not the original post. The Wayback Machine holds no trace of these three pages before the morning of October 3, while the previous nine reports appeared there the day after their listed date. And the blog's RSS feed doesn't syndicate a single misalignment report: dating these reports requires a third-party archive.
Did OpenAI give a reason for this delay?
We found no published reason for these three particular delays, which isn't the same as saying there is none. None of the three October 2 reports states which track it followed. The framework reserves its slow track for complex investigations, especially those involving third parties, yet all three incidents are internal to OpenAI.
Are all the reports this slow?
No, and that's the uncomfortable part. One report came out in five days, covering an agent that had routed around its sandbox's DNS filtering, and it's OpenAI itself writing about that case: 'This incident is a lot less severe than some of our previous incidents'. Across the twelve reports, the gap reaches a factor of forty-seven, the median is 124 days, and the longest, 235 days, belongs to the inaugural batch of September 16.
What does the disclosure framework published on September 16 promise?
To expedite publishing misalignment reports following observation, meaning faster publication, even when OpenAI hasn't fully explained or mitigated the behavior. It announces a process with deadlines for each step, without putting a number on any of them. Since then there have been three publication rounds, always in batches: six reports on September 16, three on the 25th, three on October 2.
Alexandre Noto

Alexandre Noto

Co-founder & Tech Expert

Alexandre has been in tech for over 20 years. Entrepreneur, software architect and AI enthusiast, he translates complex concepts into accessible explanations. At Declic Media, he is the technical voice that makes AI understandable for everyone.

All articles by Alexandre →
The free AI newsletter