The honest answer to "which open-weight LLM should I run locally?" in October 2026 is that it is a hardware question, and the models that top the leaderboard are not the answer. The current open-weight leaders — DeepSeek's V4 line, Qwen3-Coder-480B, the GLM-5 family, Moonshot's Kimi K2 — are trillion-ish-parameter Mixture-of-Experts models. They are open in the license sense and closed in the practical sense: you can download the weights, you just can't fit them on anything you own. What people actually run at home is a much smaller, remarkably stable set.
Here is that set, by the hardware you have:
- One 16GB GPU, or a laptop → gpt-oss-20b. OpenAI's Apache-2.0 reasoning model, 21B total but only ~3.6B active per token, lands around 16GB and "just runs."
- One 24GB GPU (RTX 3090/4090) → Qwen3-30B-A3B-Instruct-2507. The community workhorse: 30.5B total, ~3.3B active, Apache-2.0, 256K context, ~18–20GB at Q4. For code, its sibling Qwen3-Coder-30B-A3B.
- A tight coding box → Devstral Small 24B. Mistral's Apache-2.0 agentic software-engineering model, ~14GB, purpose-built to be a coding agent.
- Want multimodal → Gemma 3 27B. Image-and-text, 128K context, huge quant support (note Gemma's license terms, which aren't fully OSI-free).
- One 80GB GPU, or a 64GB+ Mac → gpt-oss-120b. 117B total / ~5.1B active, Apache-2.0 — near the ceiling of what "local" means.
That's the whole runnable tier for most people. The rest of this guide is why those are the names, and the three choices — quant, tool, hardware — that decide whether any of them actually fit. If you want the capability ranking instead of the runnable one, we keep a separate open-weight leaderboard; if you specifically want coding weights and their VRAM, see open-source LLMs for coding.
The one idea: the runnable tier is stable even as the frontier balloons#
Watch the open-weight world for a year and you notice something the leaderboards hide. The frontier open weights keep getting bigger — every few months another lab ships a larger trillion-class MoE — but the set of models you can actually load on a desktop barely moves. It sits in the same two shapes it has for a while now:
- Small dense models — roughly 4B to 32B parameters, every one active on every token (Gemma 3, the Qwen3 dense ladder, Devstral).
- MoE models with a tiny active-parameter count — big on disk, small in compute, because only a few billion parameters fire per token (gpt-oss, Qwen3-30B-A3B).
The frontier open-weight models get the headlines. The low-active-parameter MoEs get loaded.
That second shape is the local cheat code, and it's worth understanding because it's why a "30B" model runs fine on a card that shouldn't hold it. A 30B-A3B model stores 30B parameters but only activates about 3B per token. The weights can even spill from VRAM into system RAM and the thing stays usably fast, because the per-token compute is tiny. It's the reason the single most-recommended local model is a 30B that behaves, on your GPU, like a much smaller one.
Choice 1: quantization — Q4_K_M is the default, and that's correct#
You almost never run these at full precision locally. You run a GGUF quant, and the community has converged hard on one default: Q4_K_M. It's the sweet spot — roughly 4 bits per weight, a large size cut for a small, usually-imperceptible quality cost. The rules of thumb:
- Q4_K_M — the default. Use it unless you have a reason not to.
- Q5_K_M / Q6_K — if you have VRAM headroom and want to claw back a little quality.
- Q8 — near-lossless, for when you have the memory and want it.
- Below Q3 — quality falls off a cliff; only for when it's this or nothing (IQ-quants like IQ3/IQ2 soften the fall).
One upgrade worth knowing: Unsloth's "Dynamic" GGUFs (the UD-Q4_K_XL and friends) quantize different layers to different bit-widths and generally deliver better quality at about the same file size as a plain Q4_K_M. Unsloth and bartowski are the GGUF uploaders most people trust on Hugging Face.
Choice 2: the tool — Ollama, LM Studio, or llama.cpp#
Three tools cover almost everyone, and the split is about how much you want to touch:
- Ollama — easiest. A
llama.cppbackend with excellent defaults;ollama run qwen3and you're talking to a model. The common self-host stack is Ollama plus Open WebUI. - LM Studio — a polished GUI, supports both GGUF and Apple's MLX format, and is the usual recommendation for Mac users and beginners.
- llama.cpp — the engine under the other two. Run it directly for maximum control and for day-one support of brand-new model architectures, which the wrappers sometimes lag on.
For multi-user throughput or multi-GPU serving you'd reach for vLLM or SGLang, but that's a server story, not a single-desktop one. We compare the day-to-day options in Ollama vs LM Studio vs llama.cpp.
Choice 3: the hardware — and why a Mac quietly became the way to run the big ones#
For a single GPU, the number that matters is VRAM, and the Q4_K_M math is simple enough to keep in your head:
- ~8B → ~6GB (an 8GB card)
- ~14B → ~10GB (a 12GB card)
- ~27–32B → ~18–20GB (a 24GB card — the 3090/4090 sweet spot; see the cheapest 16GB card for the entry point)
- ~70B → ~40GB (two 24GB cards, or one 48GB)
Add headroom for the KV cache, which grows with context length — long context is the hidden VRAM cost, and trimming the context window is the first thing to try when a model won't load.
The quieter shift is Apple Silicon. A Mac's unified memory — up to 512GB on a Mac Studio — is addressable by the GPU, so a Mac can load 70B–120B-class models that no consumer NVIDIA card can hold, and MoEs run well on it. The trade-offs are real: prompt processing (prefill) is slower than on NVIDIA, and you won't be fine-tuning with CUDA. But for running the biggest models the local world offers, a maxed Mac has become the budget path, and it's why "what can I run at home" now has a different answer depending on whether you bought a GPU or a Mac.
The bottom line#
The leaderboard and your GPU are answering two different questions. The leaderboard asks what is the most capable open model, and the answer is a trillion-parameter MoE you'll rent. Your GPU asks what can I load tonight, and the answer has been steady for a while: a low-active-parameter MoE or a small dense model, quantized to Q4_K_M, served by Ollama or LM Studio. Start with Qwen3-30B-A3B on a 24GB card or gpt-oss-20b on 16GB, reach for a Mac's unified memory when you want the 120B class, and ignore the top of the leaderboard until you're renting a server — because that's the only place those models were ever going to run.



