Kimi K3 Ties Claude Opus 4.8 at 40% of the Price

6 min read
Article

A Chinese model lands within a single point of the Western frontier on an independent leaderboard, at 40% of the price. The gap between frontier labs is changing shape.

The free AI newsletter
Kimi K3 Ties Claude Opus 4.8 at 40% of the Price

57 to 56. A one-point gap on a leaderboard nobody at OpenAI or Anthropic seems eager to talk about this week. The 57 belongs to Kimi K3, the new model out of Chinese lab Moonshot AI. The 56 is Claude Opus 4.8, marketed for months as the best model money can buy. And that number didn't come from a Moonshot press release. It came from an independent scorekeeper.

The Judge Isn't the Seller

The leaderboard in question is Artificial Analysis's Intelligence Index, an aggregator that runs every model through the same battery of tests and publishes the results. On that index, Kimi K3 lands 4th out of 189 models tested. Ahead of it: only Claude Fable 5 and GPT-5.6 Sol. Right behind it, by a single point: Opus 4.8 and GPT-5.5.

That detail changes everything. Until now, Chinese labs mostly shone on their own benchmarks, the kind they pick and frame themselves. We've covered that playbook before, in the benchmark war between Grok and GPT, where every player releases the chart that flatters it. This time it's a different exercise. The judge isn't the seller. A third party measured blind, and it just put a Chinese model on par with the best product the West has to offer.

A Giant That Only Opens a Few Shelves

Kimi K3 is a technical beast: 2.8 trillion parameters, which would make it the largest model ever built for open weights. But the headline number is misleading. It's a mixture-of-experts architecture: out of 896 specialized subnetworks, only 16 activate to process any given token, less than 2% of the total. The model has the footprint of a national library, but for any given question it only opens a few shelves. The real compute cost stays far below what the parameter count implies.

That engineering detail has a direct effect on the bill, and that's where things get serious. Moonshot prices Kimi K3 at $3 per million input tokens and $15 per million output tokens. Opus 4.8, in its standard tier, runs $5 and $25. That's 40% cheaper, on the nose. On a full task, as measured by Artificial Analysis, the gap widens further: $0.94 for K3 versus $1.80 for Opus 4.8, nearly half.

There's context here that makes the result even more striking. Moonshot built this model under US restrictions on the most advanced chips. A Chinese lab is matching the Western frontier with, in theory, tighter access to the hardware that makes training possible in the first place. The hardware-advantage argument, long waved around as a moat, is taking on water in two places at once: price and access to silicon.

What This Gap Is Starting to Crack

For two years, frontier labs have sold a simple story: their edge is capability. Models nobody else knows how to build, which justify premium pricing and ten-figure funding rounds. That entire economy rests on one implicit bet: capability is scarce, so it's expensive, so it's defensible.

Kimi K3 cracks that bet. If a Chinese model, soon downloadable by anyone, matches the best Western product at 40% of the price, then the gap is no longer about capability. It's about price. And a price gap doesn't hold up over time: one competitor cuts rates and it evaporates. That's exactly the tension raised by the question of AI sovereignty at a global scale. Does paying a premium for the frontier still make sense once the open alternative catches up at half the cost?

The Nuance the Headlines Keep Skipping

Two caveats, because rigor is still the job here. First, an index is still an index. The Intelligence Index is a composite score: it averages dozens of tests and smooths out the rough edges. A one-point gap between K3 and Opus 4.8 doesn't mean one is definitively "better," only that they're neck and neck on this particular panel. On other tasks, the order flips: K3 takes the top spot on the Frontend Code Arena but still trails Fable 5 and GPT-5.6 overall.

Second, "open weights" is still a promise, not a fact. As of July 20, Kimi K3 runs through Moonshot's app and API, but the weights aren't downloadable yet. The lab has promised them for July 27. Until they're actually published, the "largest open model in history" is a press release, not a file. The distinction matters: there's a world of difference between a model you can audit and run on your own servers and a service you're renting like any other.

If the Weights Really Ship, the Stakes Change Scale

And that's where things get genuinely interesting. If the weights land on the 27th as promised, and under a license that actually permits real-world use, this stops being a leaderboard story. A model at this level that you can install on your own servers is a real shift: no more per-request bill to a distant API, data that never leaves the machine, and no dependency on a vendor who can change its prices, its terms, or cut off access overnight.

For a company, a government, a country that wants AI without renting it from a foreign giant, having a frontier-level model running locally would be a rare opportunity. You go from renting a tool to owning one. Only under that one condition, weights actually published and freely usable, would Kimi K3 stop being a good score on a chart and become a strategic lever. The promise is enormous. It has a date attached: the 27th.

The Real Signal

Still, the trajectory is clear, and it deserves more than a shrug or a nervous laugh. A year ago, the story was China "catching up." Today, an independent third party places a Chinese model within a single point of the Western summit, at a price that makes the premium hard to justify, and soon with weights anyone will be able to install themselves.

This fits a broader pattern, one where AI tooling is splitting into two stacks, one Western and one Chinese, each building out its own software layer. A cheap, open model accelerates that split: it hands the rest of the world a credible alternative to paying frontier prices.

The debate isn't whether the capability gap is real anymore. It's how long customers will keep paying double for a single point on an index. Moonshot just put that number on the table, and it's a stubborn one: 57 to 56.

Topics covered:

GeopoliticsAnthropicAnalysis

Frequently asked questions

Is Kimi K3 really better than Claude Opus 4.8?
On Artificial Analysis's Intelligence Index, an independent leaderboard, Kimi K3 scores 57 to Opus 4.8's 56. A one-point gap: the two models are neck and neck, and neither one is definitively better.
Is Kimi K3 open source?
Not yet. Moonshot has announced open weights for July 27, 2026, but as of July 20 they aren't downloadable and the license is still unknown. Open weights, by the way, isn't the same thing as open source.
How much does Kimi K3 cost compared to Opus 4.8?
Moonshot prices Kimi K3 at $3 per million input tokens and $15 per million output tokens, versus $5 and $25 for Opus 4.8, a 40% discount. On a full task measured by Artificial Analysis, the gap widens to nearly half the cost.
Who builds Kimi K3?
Kimi K3 comes from Moonshot AI, a Chinese lab operating under US restrictions on access to the most advanced chips.
The free AI newsletter