If you're picking an open-weight coding model for a product, the benchmark is the wrong first question. The two questions that actually decide the build are: can I legally ship it, and can I physically run it. The license answers the first. The VRAM answers the second. And here's the part most roundups skip — the model that tops the coding board is almost never the one that clears both.

So before any leaderboard, three facts a founder needs:

  1. "Smartest" and "shippable" are different lists. The strongest open coders — DeepSeek's V3.2/V4 line, Qwen3-Coder-480B, Kimi K2, GLM's larger checkpoints — are trillion-ish-parameter giants. The ones you can own, run on your own GPU, and legally build on are a shorter, different set.
  2. "Open weights" is not "open license." You can download all of these. You cannot build a product on all of them under the same terms. Apache-2.0 and MIT you ship freely; "modified" and custom licenses carry conditions you have to read first.
  3. License can change per checkpoint. The same model family can ship one weight under Apache-2.0 and the next under a custom license. The family name doesn't tell you — the model card does.

The license map (read this before the benchmark)#

For a founder, this is the ranking that decides your business. Grouped by what you're actually allowed to do:

The non-obvious trap is the last two bullets colliding: a model family you've vetted as MIT can release its next flagship under a custom license, and a "modified" MIT reads like MIT until the one clause that isn't. Verify the license on the exact checkpoint you download, every time — not the family, not last quarter's card.

The list that matters if you're running it yourself#

This is the answer to "which open coding model fits my GPU." Every figure is approximate, at Q4-class quantization, for a single 24GB consumer card (RTX 3090/4090) unless noted — and all three of the single-GPU picks are Apache-2.0, so the license question is already answered for you:

ModelType / sizeApprox. VRAM (Q4)LicenseBest for
Qwen3-Coder-30B-A3BMoE, 30.5B (3.3B active)~17–20GBApache-2.0The best coding-specialized model on one 24GB card
Devstral Small 24BDense 24B~15GBApache-2.0Driving an agentic SWE harness (OpenHands-style)
gpt-oss-20bMoE, 20.9B (3.6B active)~16GBApache-2.0Most headroom; fits a 16GB card; three reasoning levels
gpt-oss-120bMoE, 116.8B (5.1B active)one 80GB GPUApache-2.0A single H100, when 20b isn't enough
Qwen3-Coder-480B-A35BMoE, 480B (35B active)datacenterApache-2.0The best open coder you can own — but you rent the hardware

Two practical notes that trip people up. MoE memory: a model like Qwen3-Coder-30B-A3B activates only ~3B parameters per token, which makes it fast, but all 30B weights still have to live in VRAM — so budget for the total, not the active count. This is exactly why the reported 2026 "ultra-sparse" coder Qwen3-Coder-Next (~80B total, ~3B active) does not fit a 24GB card despite its tiny active count: ~40–45GB of weights have to be resident, so it wants a 48GB card or two 24GB GPUs. Quantization: Q4_K_M is the standard trade — roughly a 75% size cut for little quality loss — and Unsloth's dynamic GGUFs give slightly better quality at the same size; step up to Q5/Q6 if you have the VRAM, because low-bit quantization bites coding accuracy harder than it bites chat.

So which one should you actually pick?#

The leaderboard is a good place to start and a terrible place to stop. The coder at the top of the board this month may be a trillion parameters you'll never run on hardware you own, under a license whose next checkpoint you'll have to re-read. The model that quietly wins your build is the one that fits your GPU, carries a license you can ship, and is good enough for the 90% of code that doesn't need the frontier. In October 2026, for most founders, that model is Apache-2.0 and fits in 24GB.