Moonshot dropped Kimi K3's benchmark card in the same week its open weights land — the API went live July 16, the full 2.8-trillion-parameter weights by July 27. The scores answer the only question a founder actually has about a new model: not "is it good," but "is it good at the thing I'm about to point it at."

Here is the short version, front-loaded so you can cite it and move on: **K3 is the best open-weight model ever shipped at sustained agentic execution. It is not the best at *deep one-shot reasoning.*** Those are different jobs, and the card separates them cleanly.

Where K3 wins outright#

The one number to internalize: SWE Marathon 42.0 vs 35.0. Most benchmarks reward cracking one hard prompt. SWE Marathon rewards staying coherent across forty steps and a dozen files — which is what an autonomous coding agent actually does. K3 leads there.

Where it loses#

Across roughly 14 shared benchmarks, Fable 5 wins about 8 and K3 about 6. That's close — a trade of blows, not a blowout — and the wins are not randomly distributed. Fable 5 takes the deep-reasoning tests; K3 takes the sustained-execution and frontend ones. (If those two are your finalists, we turned this into a straight buy decision in Kimi K3 vs Claude Fable 5.)

The routing decision that falls out of the card#

You don't pick one model. You route.

This tracks the architecture, which is why it's likely to hold rather than being a benchmarking fluke. K3 runs an always-on thinking mode and Kimi Delta Attention over a 1M-token context — a design tuned to hold a long plan and a big working set stable across many tool calls. Sustained execution is what that architecture is for.

The caveat that keeps you honest#

Read the provenance. The coding scores come largely from Moonshot's own card; the Frontend Code Arena rank and the Artificial Analysis Intelligence Index placement (4th of 189, score 57 — level with Opus 4.8 and GPT-5.5) are third-party. Vendor cards select flattering benchmarks. Treat the wins as directionally real, then re-run the two or three benchmarks closest to your workload before you commit a routing rule to them. A 5-point gap on someone else's harness can invert on yours.

But the shape is trustworthy, because it's the same shape three different lenses show: an open model that has caught the closed frontier on execution and frontend, and hasn't quite caught it on the deepest reasoning. For most founders shipping agents, that's not a gap that matters — it's a discount that does.

If you're weighing whether to run it yourself: the weights are ~1.4TB and the API is almost always the right answer — we did the rent-vs-self-host math here.