Local AI · Recipe · 1× Mac

Qwen3.8-27B with MLX on 1× Mac

Qwen3.8-27B with MLX on 1× Mac: 53.3 tok/s according to github.com (dataset as of Sep 29, 2026).

by Weschera

53.3tok/sEveryday

1 request, prose prompt, with MTP (nativ) (per recipe)

Source: github.com
Intelligence (original model) · Artificial Analysisno independent value

Engine

Engine
oMLX 0.6.3rc2
Quantization
oQ4e (4-bit affine g64, 166 sens. Tensoren 5-bit; Stock-Qwen3.8-27B-Konversion)
Model family
Qwen3.8-27B
Context
262,144
Parameters
27B (dicht, lt. Modellname)
Creator
Weschera
GitHub stars
64
Repo updated
Aug 21, 2026

Measurements

Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.

Decode · 1 request

  • Everyday (decode)53.3 tok/sMittel aus 3 Reps, temp 0, thinking off, 320 gen. Tokens, ~2,1–2,4K-Token-Prompts; +11 % vs. 48.0 (oMLX 0.6.1) Source 
  • Peak (decode)72.1 tok/sMTP-Acceptance Code 99,6 % (tok/cycle 3.76); Prose-Acceptance 78 % Source 

Prefill by context

  • 4K274 tok/sANE-Prefill (vorher ~83); raw JSON ane-mtp3-results.json im Repo Source 

What you need

Hardware
1 × Mac / Apple Silicon
Weights
Engine
oMLX 0.6.3rc2
Context
262,144 tokens

Related recipes

← Back to overview

Local AI in your company?

In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.