If you're picking an open-weight coding model for a product, the benchmark is the wrong first question. The two questions that actually decide the build are: can I legally ship it, and can I physically run it. The license answers the first. The VRAM answers the second. And here's the part most roundups skip — the model that tops the coding board is almost never the one that clears both.
So before any leaderboard, three facts a founder needs:
- "Smartest" and "shippable" are different lists. The strongest open coders — DeepSeek's V3.2/V4 line, Qwen3-Coder-480B, Kimi K2, GLM's larger checkpoints — are trillion-ish-parameter giants. The ones you can own, run on your own GPU, and legally build on are a shorter, different set.
- "Open weights" is not "open license." You can download all of these. You cannot build a product on all of them under the same terms. Apache-2.0 and MIT you ship freely; "modified" and custom licenses carry conditions you have to read first.
- License can change per checkpoint. The same model family can ship one weight under Apache-2.0 and the next under a custom license. The family name doesn't tell you — the model card does.
The license map (read this before the benchmark)#
For a founder, this is the ranking that decides your business. Grouped by what you're actually allowed to do:
- Ship it freely — Apache-2.0: the Qwen3-Coder family (480B-A35B and the 30B-A3B local build), Devstral Small (Mistral's open agentic-SWE model), and OpenAI's gpt-oss-120b / 20b. Host it, fine-tune it, sell what you build — minimal strings.
- Ship it freely — MIT: DeepSeek (the V3.2 weights, and the V4 line carries MIT on its public cards) and GLM-4.6 from Z.ai. Clean and permissive.
- Read the conditions first: Kimi K2 (Moonshot) ships under a Modified MIT license that is permissive until your product crosses large monthly-active-user or revenue thresholds, at which point it requires prominent "Kimi K2" attribution in your interface. That's a clause you want to know about before, not after, you scale.
The non-obvious trap is the last two bullets colliding: a model family you've vetted as MIT can release its next flagship under a custom license, and a "modified" MIT reads like MIT until the one clause that isn't. Verify the license on the exact checkpoint you download, every time — not the family, not last quarter's card.
The list that matters if you're running it yourself#
This is the answer to "which open coding model fits my GPU." Every figure is approximate, at Q4-class quantization, for a single 24GB consumer card (RTX 3090/4090) unless noted — and all three of the single-GPU picks are Apache-2.0, so the license question is already answered for you:
| Model | Type / size | Approx. VRAM (Q4) | License | Best for |
|---|---|---|---|---|
| Qwen3-Coder-30B-A3B | MoE, 30.5B (3.3B active) | ~17–20GB | Apache-2.0 | The best coding-specialized model on one 24GB card |
| Devstral Small 24B | Dense 24B | ~15GB | Apache-2.0 | Driving an agentic SWE harness (OpenHands-style) |
| gpt-oss-20b | MoE, 20.9B (3.6B active) | ~16GB | Apache-2.0 | Most headroom; fits a 16GB card; three reasoning levels |
| gpt-oss-120b | MoE, 116.8B (5.1B active) | one 80GB GPU | Apache-2.0 | A single H100, when 20b isn't enough |
| Qwen3-Coder-480B-A35B | MoE, 480B (35B active) | datacenter | Apache-2.0 | The best open coder you can own — but you rent the hardware |
Two practical notes that trip people up. MoE memory: a model like Qwen3-Coder-30B-A3B activates only ~3B parameters per token, which makes it fast, but all 30B weights still have to live in VRAM — so budget for the total, not the active count. This is exactly why the reported 2026 "ultra-sparse" coder Qwen3-Coder-Next (~80B total, ~3B active) does not fit a 24GB card despite its tiny active count: ~40–45GB of weights have to be resident, so it wants a 48GB card or two 24GB GPUs. Quantization: Q4_K_M is the standard trade — roughly a 75% size cut for little quality loss — and Unsloth's dynamic GGUFs give slightly better quality at the same size; step up to Q5/Q6 if you have the VRAM, because low-bit quantization bites coding accuracy harder than it bites chat.
So which one should you actually pick?#
- You want the smartest open coder and you'll call it over an API: DeepSeek's V4 line, Qwen3-Coder-480B, or Kimi — but price the hosted API against owning the hardware, because unless you keep rented GPUs busy around the clock, the API wins. The math is the same one in our GPU rental price map.
- You want to run it yourself on one GPU and ship it in a product: Qwen3-Coder-30B-A3B, Devstral Small, or gpt-oss-20b. All Apache-2.0, all fit a single card, all do real coding work.
- You want the capability ranking, not the ship-and-run one: that's a different question with its own answer — we ranked it in Open-Source LLMs for Coding, September 2026 and mapped the runnable local tier in The Open-Weight LLMs People Actually Run Locally.
- You need current benchmark numbers before you trust any of this: pull them from the live boards — the Aider polyglot leaderboard for edit accuracy, llm-stats.com and Artificial Analysis for SWE-bench Verified — and check which benchmark a claim cites, because SWE-bench Verified, SWE-bench Pro, and Terminal-Bench 2.x are not the same test.
The leaderboard is a good place to start and a terrible place to stop. The coder at the top of the board this month may be a trillion parameters you'll never run on hardware you own, under a license whose next checkpoint you'll have to re-read. The model that quietly wins your build is the one that fits your GPU, carries a license you can ship, and is good enough for the 90% of code that doesn't need the frontier. In October 2026, for most founders, that model is Apache-2.0 and fits in 24GB.



