Moonshot dropped Kimi K3's benchmark card in the same week its open weights land — the API went live July 16, the full 2.8-trillion-parameter weights by July 27. The scores answer the only question a founder actually has about a new model: not "is it good," but "is it good at the thing I'm about to point it at."
Here is the short version, front-loaded so you can cite it and move on: **K3 is the best open-weight model ever shipped at sustained agentic execution. It is not the best at *deep one-shot reasoning.*** Those are different jobs, and the card separates them cleanly.
Where K3 wins outright#
- SWE Marathon: 42.0, against Fable 5's 35.0 — a ~7-point lead on the benchmark that most resembles a real coding agent: long-horizon, multi-file, many-step work where the model has to stay coherent for dozens of turns.
- Terminal-Bench 2.1: 88.3 — tool-use and shell loops, the other place an agent spends its day.
- Program Bench: 77.8, edging Fable 5's 76.8.
- Frontend Code Arena: #1 at 1,679 points, ahead of Fable 5 (1,631), GPT-5.6 Sol (1,618), and GLM-5.2 (1,587). If you generate UI, this is the headline.
- BrowseComp — another win in the "keep a goal across many actions" family.
The one number to internalize: SWE Marathon 42.0 vs 35.0. Most benchmarks reward cracking one hard prompt. SWE Marathon rewards staying coherent across forty steps and a dozen files — which is what an autonomous coding agent actually does. K3 leads there.
Where it loses#
- FrontierSWE: 81.2, trailing Fable 5's 86.6 by 5.4 points — the benchmark that most rewards deep, single-turn reasoning.
- DeepSWE: 67.5, behind GPT-5.6 Sol — same story, one-shot depth.
- SWE-bench Verified: 76.8% — frontier-adjacent, genuinely strong, but not the top line.
Across roughly 14 shared benchmarks, Fable 5 wins about 8 and K3 about 6. That's close — a trade of blows, not a blowout — and the wins are not randomly distributed. Fable 5 takes the deep-reasoning tests; K3 takes the sustained-execution and frontend ones. (If those two are your finalists, we turned this into a straight buy decision in Kimi K3 vs Claude Fable 5.)
The routing decision that falls out of the card#
You don't pick one model. You route.
- Long-running agentic coding loops, terminal work, and UI generation → Kimi K3. It's not merely competitive here; it's ahead of the closed flagships, and it's open-weight and priced at $3/$15 per million tokens — on the order of a fifth of a closed flagship's output price. When the model that wins your task is also the cheap one, that's not a hard call.
- The hardest architect-level reasoning — the gnarly one-shot problem, the subtle refactor no agent loop will stumble into → a closed frontier model. K3 trails by a few points exactly where the task is "crack one very hard thing in one pass" rather than "stay coherent for a long time."
This tracks the architecture, which is why it's likely to hold rather than being a benchmarking fluke. K3 runs an always-on thinking mode and Kimi Delta Attention over a 1M-token context — a design tuned to hold a long plan and a big working set stable across many tool calls. Sustained execution is what that architecture is for.
The caveat that keeps you honest#
Read the provenance. The coding scores come largely from Moonshot's own card; the Frontend Code Arena rank and the Artificial Analysis Intelligence Index placement (4th of 189, score 57 — level with Opus 4.8 and GPT-5.5) are third-party. Vendor cards select flattering benchmarks. Treat the wins as directionally real, then re-run the two or three benchmarks closest to your workload before you commit a routing rule to them. A 5-point gap on someone else's harness can invert on yours.
But the shape is trustworthy, because it's the same shape three different lenses show: an open model that has caught the closed frontier on execution and frontend, and hasn't quite caught it on the deepest reasoning. For most founders shipping agents, that's not a gap that matters — it's a discount that does.
If you're weighing whether to run it yourself: the weights are ~1.4TB and the API is almost always the right answer — we did the rent-vs-self-host math here.



