We let AIs run a company. Almost all of them sank it.
Princeton handed a simulated company to 13 AIs for 500 days. Three end in the black. A rules-based script with no AI beats almost all of them.

One million dollars in the bank, zero customers, 500 days to turn it into a profitable company. That's the sandbox Princeton researchers handed to thirteen AI models, with a single instruction: run the business. By the end, only three finish with more money than they started with. And one arithmetic rule, with no AI at all, beats ten of the thirteen.
A CEO in a sandbox
The test is called CEO-Bench, published in late June by Zhuang Liu's lab at Princeton. Each model runs NovaMind, a fully simulated subscription software company. It sets prices, allocates the ad budget, weighs product quality against R&D, sizes the infrastructure, handles support and negotiates with big accounts.
Thirty-four tools, nineteen data tables, 500 simulated days. The final grade is the cash left in the bank at the end. Nothing abstract here: these are the calls a real executive makes, spread across a year and a half.
And no lucky break to save the day: each model replays the game three times, and only its best run counts. The numbers below are already the high end for every AI.
The dumb rule that humiliates ten models
Here's the part that stings. The researchers slipped a brainless competitor into the race: a script that follows fixed rules. Frozen prices, frozen quotas, focus on a few customer segments, capacity tuned to recent demand. No adaptation, no strategy. Just discipline.
That script finishes fourth, with $15.7 million in the bank. It beats ten of the thirteen models tested. Five of them went flat-out bankrupt, account at zero. Others survived while destroying value: Qwen 3.7 Max and Claude Opus 4.7 land around $400,000, less than half their starting stake.
Why does the dumb rule win? Because it never reconsiders. The models do: they change prices, launch a campaign, backpedal, and burn cash with every hesitation.
Three models go the distance
Three AIs do pull ahead, and by a wide margin. Claude Fable 5 ends at $47 million, Claude Opus 4.8 at $27.8 million, GPT-5.5 at $21.3 million. So AI can run a company. The catch is that it rarely does so reliably.
GPT-5.5 sums it up. Across its three runs, it sank twice. Only one model, Fable 5, stays in the black over several attempts. The rest string together one good run and three crashes, or crash from the start.
What the benchmark really measures
Running a company isn't about nailing one move. It's about holding a coherent strategy across hundreds of chained decisions, without contradicting yourself or panicking when the numbers dip. That's exactly where most models break: they know how to open a game brilliantly, far less how to close it.
The myth of AI replacing executives runs into a mundane obstacle, then. The real bar to clear, today, is that fixed-rule script. Over 500 days, its dumb consistency is enough to beat ten of thirteen frontier models.



