1 request, prose prompt, with MTP Source
Local AI · Recipe · 1× Mac
Qwen3.8-27B with MLX on 1× Mac
Qwen3.8-27B with MLX on 1× Mac: 53.3 tok/s according to github.com (dataset as of Sep 29, 2026).
by Weschera
Engine
- Engine
- oMLX 0.6.3rc2
- Quantization
- oQ4e (4-bit affine g64, 166 sens. Tensoren 5-bit; Stock-Qwen3.8-27B-Konversion)
- Model family
- Qwen3.8-27B
- Context
- 262,144
- Parameters
- 27B (dicht, lt. Modellname)
- Creator
- Weschera
- GitHub stars
- 64
- Repo updated
- Aug 21, 2026
Measurements
Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.
Decode · 1 request
Prefill by context
What you need
- Hardware
- 1 × Mac / Apple Silicon
- Weights
- Main weightsJundot/Qwen3.8-27B-oQ4e-mtp License unclear
- Engine
- oMLX 0.6.3rc2
- Context
- 262,144 tokens
Related recipes
1 request, realistic prompt, with Qwen MTP3 probabilistic mit standard rejection (per recipe) Source
- Custom kernel required
1 request, prose prompt, with MTP (per recipe) Source
- Experimental
- Custom kernel required
1 request, mixed prompt set, context 150000, with DFlash2 (per recipe) Source
- Custom kernel required
1 request, mixed prompt set, with MTP Source
- License unclear
1 request, mixed prompt set, with DFlash2 (per recipe) Source
- Custom kernel required
Local AI in your company?
In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.