Local AI · Recipe · 1× Mac

Bonsai-2-27B Q8_0 with llama.cpp on 1× Mac

Bonsai-2-27B Q8_0 with llama.cpp on 1× Mac: 10.6 tok/s according to github.com (dataset as of Sep 29, 2026).

by Weschera

10.6tok/sEveryday

1 request, prose prompt, no speculative decoding

Source: github.com
Intelligence (original model) · Artificial Analysisno independent value

Engine

Engine
llama.cpp (prism fork)
Quantization
Ternary PQ2_0 (empfohlen) / PTQ1_0 (5,9 GB); mmproj-Q8_0 für Vision
Model family
Bonsai-2-27B (Qwen3.8-27B-Basis)
Context
131,072
Parameters
27B ternär 1.72 bits/weight (README); PQ2_0 7,2 GB
Creator
Weschera
GitHub stars
0
Repo updated
Sep 18, 2026

Measurements

Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.

Decode · 1 request

  • Everyday (decode)10.6 tok/sMac mini M4 24 GB, PQ2_0 — 3-Rep-Ladder WÄHREND parallel laufendem Spark-Bench (GPU-Contention); idle-Qualifizierung 11.0 → approx; TTFTs dieser Ladder sind Queueing-Artifacts Source 
  • Peak (decode)34.5 tok/sPQ2_0 Studio; 3 Reps Mittel, temp 0, thinking off, 320 tokens, ~2,1–2,4K prompts (bench-chat.py); flat prose=code → kein Spec-Decoding, rein bandgebunden Source 

Prefill by context

  • Prompt not stated242 tok/sPQ2_0 Studio (PTQ1_0: 205); Ternary-Unpacking macht Prefill langsamer als 4-bit-affin (242 vs 274) Source 

What you need

Hardware
1 × Mac / Apple Silicon
Weights
Engine
llama.cpp (prism fork)
Context
131,072 tokens

Notes

What matters before you rebuild it.

  • Custom kernel required

Sources

Related recipes

← Back to overview

Local AI in your company?

In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.