The honest answer to "which open-weight LLM should I run locally?" in October 2026 is that it is a hardware question, and the models that top the leaderboard are not the answer. The current open-weight leaders — DeepSeek's V4 line, Qwen3-Coder-480B, the GLM-5 family, Moonshot's Kimi K2 — are trillion-ish-parameter Mixture-of-Experts models. They are open in the license sense and closed in the practical sense: you can download the weights, you just can't fit them on anything you own. What people actually run at home is a much smaller, remarkably stable set.

Here is that set, by the hardware you have:

That's the whole runnable tier for most people. The rest of this guide is why those are the names, and the three choices — quant, tool, hardware — that decide whether any of them actually fit. If you want the capability ranking instead of the runnable one, we keep a separate open-weight leaderboard; if you specifically want coding weights and their VRAM, see open-source LLMs for coding.

The one idea: the runnable tier is stable even as the frontier balloons#

Watch the open-weight world for a year and you notice something the leaderboards hide. The frontier open weights keep getting bigger — every few months another lab ships a larger trillion-class MoE — but the set of models you can actually load on a desktop barely moves. It sits in the same two shapes it has for a while now:

  1. Small dense models — roughly 4B to 32B parameters, every one active on every token (Gemma 3, the Qwen3 dense ladder, Devstral).
  2. MoE models with a tiny active-parameter count — big on disk, small in compute, because only a few billion parameters fire per token (gpt-oss, Qwen3-30B-A3B).

The frontier open-weight models get the headlines. The low-active-parameter MoEs get loaded.

That second shape is the local cheat code, and it's worth understanding because it's why a "30B" model runs fine on a card that shouldn't hold it. A 30B-A3B model stores 30B parameters but only activates about 3B per token. The weights can even spill from VRAM into system RAM and the thing stays usably fast, because the per-token compute is tiny. It's the reason the single most-recommended local model is a 30B that behaves, on your GPU, like a much smaller one.

Choice 1: quantization — Q4_K_M is the default, and that's correct#

You almost never run these at full precision locally. You run a GGUF quant, and the community has converged hard on one default: Q4_K_M. It's the sweet spot — roughly 4 bits per weight, a large size cut for a small, usually-imperceptible quality cost. The rules of thumb:

One upgrade worth knowing: Unsloth's "Dynamic" GGUFs (the UD-Q4_K_XL and friends) quantize different layers to different bit-widths and generally deliver better quality at about the same file size as a plain Q4_K_M. Unsloth and bartowski are the GGUF uploaders most people trust on Hugging Face.

Choice 2: the tool — Ollama, LM Studio, or llama.cpp#

Three tools cover almost everyone, and the split is about how much you want to touch:

For multi-user throughput or multi-GPU serving you'd reach for vLLM or SGLang, but that's a server story, not a single-desktop one. We compare the day-to-day options in Ollama vs LM Studio vs llama.cpp.

Choice 3: the hardware — and why a Mac quietly became the way to run the big ones#

For a single GPU, the number that matters is VRAM, and the Q4_K_M math is simple enough to keep in your head:

Add headroom for the KV cache, which grows with context length — long context is the hidden VRAM cost, and trimming the context window is the first thing to try when a model won't load.

The quieter shift is Apple Silicon. A Mac's unified memory — up to 512GB on a Mac Studio — is addressable by the GPU, so a Mac can load 70B–120B-class models that no consumer NVIDIA card can hold, and MoEs run well on it. The trade-offs are real: prompt processing (prefill) is slower than on NVIDIA, and you won't be fine-tuning with CUDA. But for running the biggest models the local world offers, a maxed Mac has become the budget path, and it's why "what can I run at home" now has a different answer depending on whether you bought a GPU or a Mac.

The bottom line#

The leaderboard and your GPU are answering two different questions. The leaderboard asks what is the most capable open model, and the answer is a trillion-parameter MoE you'll rent. Your GPU asks what can I load tonight, and the answer has been steady for a while: a low-active-parameter MoE or a small dense model, quantized to Q4_K_M, served by Ollama or LM Studio. Start with Qwen3-30B-A3B on a 24GB card or gpt-oss-20b on 16GB, reach for a Mac's unified memory when you want the 120B class, and ignore the top of the leaderboard until you're renting a server — because that's the only place those models were ever going to run.