1 request, mixed prompt set, no speculative decoding Source
Local AI · Recipe · 8× RTX 3090
GLM-4.7 218B AutoRound INT4 with vLLM on 8× RTX 3090
by 0xSero
Intelligence (original model) · Artificial Analysisno independent value
Engine
- Engine
- vLLM
- Quantization
- W4A16 AutoRound (group 128) + FP8 KV
- Model family
- GLM-4.7
- Context
- 165,000
- Parameters
- 218B total / 32B active (REAP)
- Creator
- 0xSero
- GitHub stars
- 25
- Repo updated
- May 30, 2026
Measurements
Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.
No sourced measurements for this recipe.
What you need
- Hardware
- 8 × NVIDIA RTX 3090
- Weights
- Main weights0xSero/GLM-4.7-218B-W4A16 Apache-2.0
- Engine
- vLLM
- Context
- 165,000 tokens
Notes
What matters before you rebuild it.
- REAP-pruned
Sources
Related recipes
8× RTX 3090Qwen3.5-397B-A17B (REAP 264B)
62tok/sMixed
Intelligenceno independent value
8× RTX 3090
13.9tok/sPeak
1 request, best value reported by source Source
Intelligence19with thinkingno thinking 15
- Experimental
Local AI in your company?
In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.