Moonshot AI unveiled Kimi K3 on July 16, 2026, and the number in the headline is the one everyone is repeating: 2.8 trillion parameters, open weights, July 27. It is, if the weights ship as promised, the first open-weight model to reach the three-trillion-parameter class. That's a real milestone. It is also not, by itself, a decision. Here is what the model actually is, and what a founder does with it this month.
The verified specs, in one screen#
- Architecture: a 2.8-trillion-parameter mixture-of-experts model with 896 experts, 16 active per token — so any given forward pass activates a small fraction of the total.
- Quantization: MXFP4 (Microscaling FP4), 4-bit weights with per-block scaling, natively supported on NVIDIA Blackwell and AMD MI400. Full weights come to roughly 1.4TB of storage.
- Context: a 1-million-token window, with pricing flat across the whole thing.
- Access today: hosted-only, via the Kimi API and OpenRouter, at $3.00 per million input tokens and $15.00 per million output ($0.30 cached input).
- Open weights: promised by July 27, 2026 under a Modified MIT license — but at launch, no checkpoint, license file, or model card had been published.
Everything above is verifiable. The part you should hold loosely is the benchmark story.
The benchmark caveat, said plainly#
Moonshot's coding claims rest on newer suites — DeepSWE, SWE Marathon, Program Bench — that are not yet independently replayable, and there were no SWE-bench Verified or Pro figures at launch. K3 reportedly leads on sustained-coding benchmarks, which would suggest strength in long agent sessions, but "reportedly leads on a benchmark its maker highlighted" is a marketing sentence, not an evaluation. The only benchmark that decides anything for you is your own task set. Run it there.
"Open weight" and "you can build a business on it" are not the same sentence until you've read the license — and at launch, the license wasn't published.
The decision the headline hides#
The instinct when a near-frontier model goes open is to reach for GPUs. Resist it for a beat. A 1.4TB model is a multi-GPU or rented-inference deployment, not a single-card side project — the infrastructure and ops burden is real, and it only pays off past a certain token volume. For most solo founders and small teams, the sequence is:
- Prototype on the hosted API now. $3/$15 per million tokens is competitive, and you get to answer the only question that matters first: does the quality hold on your workload?
- Measure your real token volume. The case for self-hosting is a cost curve — it bends in your favor only when the per-token API bill exceeds the amortized cost of running the weights yourself.
- Read the actual license before you build on it. "Modified MIT" is a promise, not a document, until Moonshot publishes the terms. The modifications are where the constraints live.
This is not a Kimi-specific caution — it is the standing playbook for every open-weight release, and we've walked the routing and licensing tradeoffs across the field before: GLM 5.2 vs MiniMax M3 vs Kimi K2 for open-weight coding, and the open-weight licenses that actually govern what you ship. K3 is the biggest entry in that lineage; the questions it raises are the same ones, scaled up.
What it means for a founder#
The right read on Kimi K3 this week is neither hype nor dismissal. A 2.8T open-weight model with a 1M context and a public price is a genuine option — a lock-in hedge and a potential cost lever for teams with the volume and the data-residency reasons to run their own inference. But it becomes that only after two things happen: the weights actually ship on July 27 with a license you can live with, and your own benchmarks confirm the quality that the launch leaderboard asserts. Until then, the highest-leverage move is the cheapest one: get an API key, throw your real tasks at it, and let the results — not the parameter count — make the call.
Newer: Kimi K3 no longer has the open-weight week to itself. On July 19 Alibaba previewed Qwen3.8-Max, a 2.4-trillion-parameter multimodal model claiming "second only to Fable 5" — though its benchmarks and license aren't published yet. We put the two side by side in Qwen3.8-Max vs Kimi K3: which belongs in your stack?.



