When to Choose Each Option
Clear guidance based on your specific situation and needs.
Our Recommendation
There is no universal winner — the axis is quality-exactness versus headroom. IQ3_S is the rational default for the 16-GB point of 2026: AIME25 and LiveCodeBench reproduced exactly, GPQA-Diamond off by 0.51 points, and ~40 tok/s on an RTX 5060 Ti keeps the model above reading speed even at long context. The IQ2 tiers are different instruments, not worse ones: on a 12-GB card IQ2_XS is the only tier that fits with real headroom, and for batch-style work the smaller files buy more KV space and larger batches for a 3–5 point GPQA gap that rarely surfaces in short, format-like tasks. The decision rule: interactive and graded work on 16 GB → IQ3_S; 12-GB card or pure batch throughput → IQ2_S or IQ2_XS. Dense compression has a hard floor near 2.5 bpw, so these files already cover the useful range — and the .rco-allocation.txt files make every tier's budget verifiable.
- Choose IQ3_S (3.5 bpw, 11.8 GB) — the recommended quality tier when...
- You work interactively on a 16-GB card — chat, coding assistants, agent loops — and want base-level reasoning with nothing measurable lost.
- You need ~100k context while keeping full task quality; the 3–4 GB left by the 11.8 GB file covers your window.
- Your tasks are graded or reasoning-heavy (GPQA-style Q&A, multi-step code), where the 3–5 point gap of the IQ2 tiers actually shows up.
- You can afford the MTP variant (+0.35 GB) and want speculative decoding on top of the quality tier.
- Choose IQ2 tiers (2.75/2.5 bpw, 9.3/8.4 GB) — the headroom tiers when...
- You run on a 12-GB laptop GPU — IQ2_XS at 8.4 GB is the only tier that fits with ~1.3 GB of headroom.
- You process batch jobs (classification, extraction, many documents) where 3–5 GPQA points rarely matter and larger batch sizes dominate throughput.
- You want maximum KV-cache space for very long windows or several concurrent sessions on one card.
- You are prototyping and value fast load/teardown cycles over the last half percent of quality.