Kimi K3 Scored 57. The 130 Million Output Tokens Are the Catch.
Kimi K3 reached a 57 Artificial Analysis score, but its 130 million evaluation-output tokens, near-$1 task cost, delayed weights, and 64-accelerator serving guidance define the real trade.
AI-powered · Limited to 20 requests per hour

The Artificial Analysis page for Kimi K3 contains both the headline and the warning label. Kimi K3 scores 57 on the site's Intelligence Index, close to the leading proprietary systems in that comparison. It also generated roughly 130 million output tokens across the evaluation, versus a 63 million average for comparable models.
That combination is the story for me. Moonshot AI has built an unusually capable 2.8-trillion-parameter model and says its weights are coming. But Kimi K3 does not make frontier intelligence look small, cheap, or effortless. It makes the trade more visible: stronger open models can narrow the quality gap while widening the amount of inference work needed to finish a task.
Answer Snapshot
| Question | My read |
|---|---|
| What happened? | Artificial Analysis scored Kimi K3 at 57 on its Intelligence Index, with about 130 million evaluation-output tokens, 62 output tokens per second, and a weighted cost of about $0.94 per benchmark task. |
| What is Kimi K3? | Moonshot describes it as a 2.8-trillion-parameter, native-vision model with a one-million-token context window, built for coding, knowledge work, and reasoning. |
| Who benefits? | Teams that want a strong hosted model now, and researchers or infrastructure operators willing to wait for weights and handle a very large deployment later. |
| What is the catch? | The independent evaluation found high output volume, below-average output speed, and a cost per task close to the most expensive models in its comparison. |
| My thesis | Kimi K3 makes open-model capability more credible, but its real test is intelligence per completed task, not intelligence per benchmark point. |
A Frontier Score With A Long Tail
Artificial Analysis says its version 4.1 Intelligence Index combines nine evaluations spanning knowledge, science, banking, coding, long context, and terminal work. On that aggregate, Kimi K3's score of 57.1 sits behind Claude Fable 5 with fallback at 59.9 and GPT-5.6 Sol max at 58.9, while landing ahead of several other proprietary and open models in the displayed comparison.
I think that is genuinely notable. This is not a vendor chart declaring victory on a hand-picked task. It is an independent harness placing Kimi K3 in the same performance neighborhood as the strongest systems it tested. Moonshot's own Kimi K3 release post is also unusually explicit that the model still trails the most powerful proprietary models in overall performance.
But the same independent page records 130 million output tokens from the Intelligence Index run, more than twice the 63 million comparison average. Artificial Analysis also reports 62 output tokens per second against a 73-token average. I read those numbers as a warning against treating the score as the whole product. A model can be near the top on capability while still taking a long, expensive path to get there.

The Token Trail Changes The Price Story
The Kimi K3 page lists hosted pricing of $3 per million input tokens and $15 per million output tokens. Across the full Intelligence Index evaluation, Artificial Analysis says the model cost $2,690.80 to run. Its weighted cost per task was about $0.94, below GPT-5.6 Sol max in the same comparison but far above several models with lower aggregate intelligence scores.
That is why I would not describe K3 as a cheap frontier model. It may be cheaper than the two leaders on this particular benchmark mix, but it is not competing on bargain-bin inference. A TechRadar review published before K3's release reported Kimi K2.6 API pricing around $0.55 per million input tokens and $2.65 per million output tokens. The comparison is not apples-to-apples capability, but it shows how dramatically K3 changes Kimi's old low-price pitch.
The public question I find most persuasive is therefore not whether K3 is “really” frontier. In a LocalLLaMA discussion of the benchmark results, developers immediately asked how many tokens max reasoning uses and whether long thinking would erase the price advantage. That is the right skepticism. Output price, verbosity, latency, and retry rate compound each other.
“Open” Is A Delivery Schedule Today
Moonshot calls Kimi K3 the first open model in the 3-trillion-parameter class, but the release post says the full weights will arrive by July 27, 2026. It also says the technical report with more architecture, training, and evaluation detail will be released alongside them. As of the announcement I inspected, the hosted model is available through Kimi's products and API; the promised weight release is still future work.
I do not think that makes the open claim meaningless. A dated commitment is more useful than a vague promise, and Moonshot published both the code repository and model weights for Kimi K2.5. But I would keep the tense precise. Kimi K3 is available as a service now. It becomes an inspectable, downloadable open-weight release when the artifacts and license actually land.
This distinction matters to the people who benefit most from open weights. Researchers need the checkpoint and technical report to reproduce claims. Infrastructure teams need model files and serving support to estimate memory, throughput, and operating cost. Product teams considering a hosted API can start earlier, but they are evaluating Moonshot's service rather than the portability promised by the future release.

The Architecture Does Not Make It Small
Moonshot says K3 activates 16 of 896 experts through its Stable LatentMoE design and combines that sparsity with Kimi Delta Attention and Attention Residuals. The company claims these changes deliver roughly 2.5 times better scaling efficiency than Kimi K2. Those are architecture claims from the developer, not independent deployment results, so I would wait for the technical report before drawing a larger conclusion.
One operational detail is already concrete enough to reset expectations. Moonshot recommends serving Kimi K3 on supernode configurations with at least 64 accelerators. Sparse activation reduces the work performed for each token relative to using every parameter, but it does not turn a 2.8-trillion-parameter checkpoint into a workstation model. Weight storage, expert routing, high-bandwidth communication, cooling, and serving software remain part of the product.
This is where the open-weight story splits into two audiences. Large labs and inference providers may gain meaningful control over a frontier-tier system. Most individual developers will consume K3 through an API or a hosted platform even after the weights appear. “Open” expands who can inspect and deploy the model; it does not mean everyone can deploy it economically.

What I Would Measure Before Switching
If I were evaluating Kimi K3 for agentic coding or research, I would not begin with the 57-point score. I would start with a task set that has known answers, fixed tool permissions, and a review rubric. Then I would measure accepted-task rate, total output and reasoning tokens, time to completion, tool-call count, retries, and human correction time.
I would also test effort controls when they exist. Moonshot says K3 launches with maximum thinking effort by default and that low- and high-effort modes will come later. The ability to spend less reasoning on routine work may matter as much as another benchmark point, because the current independent numbers suggest K3's default path is hungry.
Finally, I would separate the hosted decision from the open-weight decision. For the API, the questions are reliability, privacy terms, rate limits, latency, and successful-task cost. For self-hosting, the questions are license, quantization quality, hardware topology, serving support, and total cost of ownership. One model name does not make those the same purchase.
My Bottom Line
Kimi K3 deserves attention because the independent result supports the core capability claim. A score of 57 puts an announced open-weight model into a performance tier that was recently the exclusive territory of closed systems. That gives developers more leverage and gives the open-model ecosystem a larger target to optimize.
The 130 million output tokens keep me from turning that result into a victory lap. Frontier capability is valuable, but capability that arrives through long reasoning traces, slower generation, and a serious deployment footprint can still be the wrong operational choice. Kimi K3 has made the summit. Now I want to see whether teams can carry the token bill, receive the promised weights, and turn that intelligence into completed work efficiently.
License
News text © 2026 Mark Huang. News text may be shared or translated for non-commercial use with attribution to https://markhuang.ai/news/kimi-k3-130-million-token-catch.
Suggested attribution: Based on "Kimi K3 Scored 57. The 130 Million Output Tokens Are the Catch." by Mark Huang, originally published at https://markhuang.ai/news/kimi-k3-130-million-token-catch.
