1 request, mixed prompt set Source
Local AI · Recipe · 8× RTX 3090
Qwen3.5-397B-A17B FP8 with vLLM on 8× RTX 3090
by 0xSero
Intelligence (original model) · Artificial Analysisno independent value
Engine
- Engine
- vLLM 0.19.0
- Quantization
- W4A16 + FP8 KV
- Model family
- Qwen3.5-397B-A17B (REAP 264B)
- Context
- —
- Parameters
- 264B (REAP)
- Creator
- 0xSero
- GitHub stars
- 11
- Repo updated
- May 30, 2026
Measurements
Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.
No sourced measurements for this recipe.
What you need
- Hardware
- 8 × NVIDIA RTX 3090
- Weights
- Main weights0xSero/Qwen3.5-264B-W4A16 Apache-2.0
- Engine
- vLLM 0.19.0
Notes
What matters before you rebuild it.
- REAP-pruned
Sources
Related recipes
8× RTX 3090
38.5tok/sPeak
Intelligenceno independent value
8× RTX 3090
13.9tok/sPeak
1 request, best value reported by source Source
Intelligence19with thinkingno thinking 15
- Experimental
Local AI in your company?
In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.