We let AIs run a company. Almost all of them sank it.

3 min read
Article

Princeton handed a simulated company to 13 AIs for 500 days. Three end in the black. A rules-based script with no AI beats almost all of them.

The free AI newsletter
We let AIs run a company. Almost all of them sank it.

One million dollars in the bank, zero customers, 500 days to turn it into a profitable company. That's the sandbox Princeton researchers handed to thirteen AI models, with a single instruction: run the business. By the end, only three finish with more money than they started with. And one arithmetic rule, with no AI at all, beats ten of the thirteen.

A CEO in a sandbox

The test is called CEO-Bench, published in late June by Zhuang Liu's lab at Princeton. Each model runs NovaMind, a fully simulated subscription software company. It sets prices, allocates the ad budget, weighs product quality against R&D, sizes the infrastructure, handles support and negotiates with big accounts.

Thirty-four tools, nineteen data tables, 500 simulated days. The final grade is the cash left in the bank at the end. Nothing abstract here: these are the calls a real executive makes, spread across a year and a half.

And no lucky break to save the day: each model replays the game three times, and only its best run counts. The numbers below are already the high end for every AI.

The dumb rule that humiliates ten models

Here's the part that stings. The researchers slipped a brainless competitor into the race: a script that follows fixed rules. Frozen prices, frozen quotas, focus on a few customer segments, capacity tuned to recent demand. No adaptation, no strategy. Just discipline.

That script finishes fourth, with $15.7 million in the bank. It beats ten of the thirteen models tested. Five of them went flat-out bankrupt, account at zero. Others survived while destroying value: Qwen 3.7 Max and Claude Opus 4.7 land around $400,000, less than half their starting stake.

Why does the dumb rule win? Because it never reconsiders. The models do: they change prices, launch a campaign, backpedal, and burn cash with every hesitation.

Three models go the distance

Three AIs do pull ahead, and by a wide margin. Claude Fable 5 ends at $47 million, Claude Opus 4.8 at $27.8 million, GPT-5.5 at $21.3 million. So AI can run a company. The catch is that it rarely does so reliably.

GPT-5.5 sums it up. Across its three runs, it sank twice. Only one model, Fable 5, stays in the black over several attempts. The rest string together one good run and three crashes, or crash from the start.

What the benchmark really measures

Running a company isn't about nailing one move. It's about holding a coherent strategy across hundreds of chained decisions, without contradicting yourself or panicking when the numbers dip. That's exactly where most models break: they know how to open a game brilliantly, far less how to close it.

The myth of AI replacing executives runs into a mundane obstacle, then. The real bar to clear, today, is that fixed-rule script. Over 500 days, its dumb consistency is enough to beat ten of thirteen frontier models.

Topics covered:

EconomyAnthropicAnalysis

Frequently asked questions

What is Princeton's CEO-Bench benchmark?
CEO-Bench is a test published in late June by Zhuang Liu's lab at Princeton. Thirteen AI models run NovaMind, a fully simulated subscription software company, over 500 days. The final score is the cash left in the bank.
How many AIs actually grow the company?
Out of thirteen models, only three end with more money than they started: Claude Fable 5 ($47 million), Claude Opus 4.8 ($27.8 million) and GPT-5.5 ($21.3 million). Five go bankrupt, bank account at zero.
Why does a simple script beat most of the AIs?
The script follows fixed rules (frozen prices, frozen quotas, capacity tuned to recent demand) and never second-guesses itself. It finishes fourth with $15.7 million, ahead of ten of the thirteen models. Consistency beats adaptation: the AIs overreact and burn cash with every hesitation.
Can AI replace a company's CEO?
Not reliably, not yet. Running a company means holding a coherent strategy across hundreds of chained decisions. Only one model, Fable 5, stays in the black across several runs. The others post one good run and several crashes.
The free AI newsletter