(01)Best all-round open-weight family and the safest default: one architecture published across 0.8B, 2B, 4B, 9B, 27B, 35B-A3B and 122B-A10B, all under Apache 2.0. The flagship Qwen3.5-122B-A10B activates 10B of 122B parameters per token and ships a native 262,144-token context extensible to roughly 1M. The small variants are among the most widely used open-weight checkpoints on Hugging Face, which means a broad ecosystem of quantizations, fine-tunes and serving recipes. Start here unless a specific workload pushes you elsewhere.
General purpose across the whole size ladder — one family from laptop to multi-GPU nodeFree weights (Apache 2.0); you pay only for the GPUs you runAI-Native
(02)Best open-weight model for long-horizon agentic and coding work. GLM-5.2 is the flagship sparse-attention MoE from Z.ai, released under a plain MIT license, and it is the first in the GLM line to sustain a solid 1M-token context rather than a nominal one. Its IndexShare attention reuses the same indexer across every four sparse layers, which, according to Z.ai's model card, cuts per-token FLOPs substantially at 1M context — that efficiency is the reason it is practical to self-host at all. Independent rankings such as Onyx's self-hosted LLM leaderboard place it among the strongest self-hostable coding models.
Long-horizon agents and coding — the strongest self-hostable coder in 2026Free weights (MIT); hosted API available via Z.ai if you do not want the clusterAI-Native
(03)Best capability-per-GPU and the easiest open-weight model to actually get into production. gpt-oss-120b has 117B parameters with 5.1B active and, thanks to MXFP4 post-training quantization of the MoE weights, runs on a single 80GB GPU (H100 or MI300X); gpt-oss-20b has 21B parameters with 3.6B active and fits in 16GB, which puts it on a workstation or a high-end laptop. Apache 2.0, fine-tunable — the 120b on a single H100 node, the 20b on consumer hardware. If your constraint is "one GPU we already own," this is the shortlist of one.
Best fit for constrained hardware — frontier-adjacent reasoning on a single GPUFree weights (Apache 2.0), no copyleft or patent strings attachedAI-Native
(04)Best open-weight frontier reasoning if you have the cluster. DeepSeek-V4-Pro is a 1.6T-parameter MoE with 49B active, DeepSeek-V4-Flash is 284B with 13B active, and both support a full one-million-token context under an MIT license. The hybrid attention design (Compressed Sparse Attention plus Heavily Compressed Attention) is what makes the long context affordable: according to DeepSeek's model card, at 1M tokens V4-Pro needs only a fraction of the per-token inference FLOPs and KV cache of its predecessor V3.2. Pro is a rental-or-cluster proposition; Flash is the variant most teams will actually self-host.
Frontier reasoning and million-token document workFree weights (MIT); serious GPU budget or a hosted provider required for ProAI-Native
(05)Best open-weight coding agent when you want a model tuned for the agent loop rather than for chat. Kimi K2.7 Code is a 1T-parameter MoE with 32B active, 384 experts (8 selected plus 1 shared per token) and a 256K context, built on K2.6 and evaluated primarily inside agent harnesses rather than on static benchmarks. The catch to plan for: it ships under a Modified MIT license, not plain MIT or Apache — read the terms before you build a commercial product on it. Note also that the widely reported "Kimi K3 weights" had still not appeared in Moonshot's official Hugging Face org as of this guide's publication.
Agentic coding — tuned for long multi-step tool-use sessionsFree weights, but under a Modified MIT license — review before commercial use
(06)Best open-weight family for on-device and edge deployment. Gemma 4 ships five sizes — E2B, E4B, 12B, 26B-A4B and 31B — in both dense and MoE flavours, with a 256K context, multimodal text-and-image input (audio on E2B, E4B and 12B) and support for 140+ languages, all under Apache 2.0. If your deployment target is a phone, a laptop or an on-prem appliance rather than a GPU cluster, this family is designed for exactly that range.
On-device and edge — phones and laptops through to single-server deploymentsFree weights (Apache 2.0)AI-Native
(07)Best choice when EU data residency and sovereignty are hard requirements rather than nice-to-haves. Mistral Small 4 is a 119B-parameter model with 6.5B activated per token, a 256K context and a per-request reasoning_effort parameter, published under Apache 2.0; Mistral Large 3 (675B) is the Apache-2.0 flagship above it. A European vendor with permissive licensing is a materially different procurement conversation for regulated DACH industries — and Mistral's push into air-gapped and sovereign deployments makes it the pragmatic default there, even where a Chinese-lab model scores higher on a leaderboard.
EU-sovereign and regulated deployments — air-gapped, on-prem, data-residency-boundFree weights (Apache 2.0); commercial support and La Plateforme availableAI-Native
(08)Best open-weight model for high-throughput million-token workloads. MiniMax-M3 is a natively multimodal MoE with roughly 428B parameters and about 23B activated, built around MiniMax Sparse Attention (MSA) — which, according to MiniMax's model card, delivers much faster prefill and decode than its predecessor at 1M context. If your bottleneck is serving very long contexts cheaply rather than topping a reasoning leaderboard, that throughput profile is the differentiator. Licensed under MiniMax's own terms, so treat the license as a review item, not a formality.
High-throughput long-context serving and multimodal inputFree weights under MiniMax model licence — review terms before commercial use