1 request, mixed prompt set, with MTP (per recipe) Source
Local AI · Recipe · 1× Strix Halo
Qwen3.6-35B-A3B Q4_K_M with llama.cpp on 1× Strix Halo
by 0xSero
Engine
- Engine
- llama.cpp
- Quantization
- GGUF Q4_K_M; Variante DYNAMIC (Mixed per-tensor)
- Model family
- Qwen3.6-35B-A3B
- Context
- —
- Parameters
- 35B total / 3B active
- Creator
- 0xSero
- GitHub stars
- 26
- Repo updated
- May 30, 2026
Measurements
Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.
No sourced measurements for this recipe.
What you need
- Hardware
- 1 × AMD Strix Halo
- Weights
- Main weights0xSero/Qwen3.6-35B-GGUF Apache-2.0
- BaseQwen/Qwen3.6-35B-A3B Apache-2.0
- Engine
- llama.cpp
Related recipes
1 request, mixed prompt set, with MTP (per recipe) Source
1 request, mixed prompt set, with MTP Source
1 request, mixed prompt set, no speculative decoding Source
1 request, prose prompt, no speculative decoding Source
- License unclear
1 request, mixed prompt set, with ndt=3 Source
- Custom kernel required
- License unclear
Local AI in your company?
In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.