---
type: "Comparison"
title: "IQ3_S vs Q4_K_M: which GGUF quantization fits 27B into 16 GB"
description: "IQ3_S (GSQ+RCO) vs Q4_K_M (classic, uniform)"
resource: "https://www.contextstudios.ai/comparisons/iq3-s-vs-q4-k-m"
language: "en"
tags: ["IQ3_S", "Q4_K_M", "GGUF quantization", "16 GB VRAM", "Qwen3.8-27B", "llama.cpp"]
generated:
  by: "process:contextstudios-md/1"
  at: "2026-10-08T20:46:23.544Z"
status: "stable"
---

# IQ3_S vs Q4_K_M: which GGUF quantization fits 27B into 16 GB

The difference is a memory plan: Q4_K_M compresses a 27B model to roughly 15 GB, filling a 16 GB card almost completely — the remaining gap to the KV cache becomes the bottleneck for context length. IQ3_S uses per-tensor search (GSQ+RCO by IST-DASLab) to reach 11.8 GB while keeping 99.8 % of the task average versus the BF16 base — a 4.6× reduction. On a 16 GB laptop, IQ3_S is the rational endpoint; Q4_K_M stays the proven default where the card offers more room.

## Detailed Comparison

| Factor | IQ3_S (GSQ+RCO) | Q4_K_M (classic, uniform) | Winner |
|--------|------|------|--------|
| File size | 11.8 GB — 4.6× smaller than the BF16 base (53.8 GB) | around 15 GB — about 3.6× smaller than the BF16 base | IQ3_S (GSQ+RCO) |
| KV-cache headroom | 3–4 GB remain on 16 GB cards — contexts up to roughly 100k tokens stay practical | under 1 GB remains — long windows hit the VRAM edge early | IQ3_S (GSQ+RCO) |
| Task quality | 99.8 % task average: AIME25 100.00, GPQA-Diamond 89.39, LiveCodeBench v6 85.71 — AIME and LCB exactly at base level | near-lossless on dense models, typically 99–100 % of the FP16 average | Tie |
| Decode speed | ~40 tok/s measured on an RTX 5060 Ti 16 GB, 552 tok/s prefill at 112k context | similarly fast on larger cards; on 16 GB the smaller file streams less weight memory | IQ3_S (GSQ+RCO) |
| Compatibility | plain GGUF file, runs out of the box in llama.cpp, Ollama and LM Studio | established default across all GGUF tooling | Tie |
| Long-context stability | 17–30 tok/s at the end of a full 112k window — the rate falls with KV-cache growth but stays usable | window and cache fight over the last gigabytes; OOM risk in the upper range | IQ3_S (GSQ+RCO) |
| Auditability | tensor allocation shipped as .rco-allocation.txt plus importance matrix (1000 × 4096 tokens) in the repo | standardised, well-documented scheme without a tensor audit file | IQ3_S (GSQ+RCO) |
| Multimodal support | identical mmproj file (BF16, 0.9 GB) for all IQ tiers; MTP build +0.35 GB | same mmproj file — no disadvantage, it is identical for both tiers | Tie |

## Key Statistics

- **Qwen3.8-27B as IQ3_S: 11.8 GB at 3.50 bpw with AIME25 100.00 / GPQA-Diamond 89.39 / LiveCodeBench v6 85.71 — versus 53.8 GB BF16 base (GPQA 89.90)** — [ISTA-DASLab model card](https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF) (2026)
- **~40 tok/s average with IQ3_S on the RTX 5060 Ti 16 GB; 552 tok/s prefill at 112k context; 25–30 tok/s near the end of the full window** — [community measurements (HF discussion)](https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF/discussions/6) (2026)
- **~136 tok/s with IQ3_S on the RTX 5090, long context** — [ISTA-DASLab model card](https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF) (2026)
- **With the old uniform quantization logic a 27B dense model needed 18–20 GB; per-tensor search pushes it under the 16 GB line** — [ISTA-DASLab model card](https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF) (2026)
- **Even the smallest tier IQ2_XS (8.4 GB, 2.50 bpw) delivers 96.67 on AIME25 — above the BF16 base in zero-shot** — [ISTA-DASLab model card](https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF) (2026)
- **Dense compression has a hard floor around 2.5 bpw: below it the task average collapses** — [ISTA-DASLab model card](https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF) (2026)

## Choose IQ3_S (GSQ+RCO) when...

- 16 GB VRAM is the hard ceiling
- contexts up to ~100k with real headroom are required
- the task average must stay above 99 %
- you want a reproducible tensor allocation with audit file

## Choose Q4_K_M (classic, uniform) when...

- more than 18 GB of VRAM is available
- the toolchain expects the uniform, standard-documented scheme
- existing Q4_K_M caches should not be rebuilt
- maximum bit width per layer with standard support matters

## Our Recommendation

IQ3_S wins on 16 GB cards in three dimensions — file size, KV headroom, long-context stability — at practically identical quality (99.8 % task average, AIME25 and LiveCodeBench exactly at base level). Q4_K_M keeps up where the card offers more than 18 GB and the uniform scheme is the expected standard. The decision rule: measure the longest context you actually use, then pick the file that leaves 3 GB for cache and activations — with that arithmetic, IQ3_S lands on 16 GB, Q4_K_M just under.

## Frequently Asked Questions

**Q: Which file do I take on a 16 GB card?**
A: IQ3_S at 11.8 GB for full model quality and 100k–112k context. IQ2_S (9.3 GB) if you need more headroom for large batches or longer windows — it already reaches base level on AIME25. IQ2_XS is the pick for 12 GB notebooks, where 8.4 GB plus mmproj still fits with room to spare.

**Q: Does that make Q4_K_M obsolete?**
A: No. On cards with more than 16 GB, Q4_K_M remains the proven near-lossless default and every toolchain understands it. The advantage of IQ3_S is a geometric consequence, not algorithmic magic: 3.4 GB less weight, 3.4 GB more cache.

**Q: Do the 40 tok/s hold for every context?**
A: Plan a span, not a number: up to ~50 tok/s early in the window, 25–30 tok/s near the end of a full 112k context. Prefill throughput stays at 552 tok/s, running in parallel over the prompt tokens.

**Q: Do I need the -mtp variant?**
A: If your harness supports speculative decoding: yes. The MTP head costs only 0.35 GB, leaves the weights unchanged and measurably speeds up decoding. Without MTP support in the harness, the standard file is enough.

