When to Choose Each Option
Clear guidance based on your specific situation and needs.
Our Recommendation
IQ3_S wins on 16 GB cards in three dimensions — file size, KV headroom, long-context stability — at practically identical quality (99.8 % task average, AIME25 and LiveCodeBench exactly at base level). Q4_K_M keeps up where the card offers more than 18 GB and the uniform scheme is the expected standard. The decision rule: measure the longest context you actually use, then pick the file that leaves 3 GB for cache and activations — with that arithmetic, IQ3_S lands on 16 GB, Q4_K_M just under.
- Choose IQ3_S (GSQ+RCO) when...
- 16 GB VRAM is the hard ceiling
- contexts up to ~100k with real headroom are required
- the task average must stay above 99 %
- you want a reproducible tensor allocation with audit file
- Choose Q4_K_M (classic, uniform) when...
- more than 18 GB of VRAM is available
- the toolchain expects the uniform, standard-documented scheme
- existing Q4_K_M caches should not be rebuilt
- maximum bit width per layer with standard support matters