---
type: "Article"
title: "Bonsai-2-27B Q8_0 with llama.cpp on 1× Mac"
description: "Bonsai-2-27B Q8_0 with llama.cpp on 1× Mac: recipe by Weschera with speed, condition and source for every number, weights and license."
resource: "https://www.contextstudios.ai/local-ai/weschera--bonsai-2-27b-pq2_0-prism-llamacpp-mac-studio-m4-max"
language: "en"
generated:
  by: "process:contextstudios-md/1"
  at: "2026-10-02T21:55:41.947Z"
status: "stable"
---

# Bonsai-2-27B Q8_0 with llama.cpp on 1× Mac

Bonsai-2-27B Q8_0 with llama.cpp on 1× Mac: 10.6 tok/s according to github.com (dataset as of Sep 29, 2026).

- Creator: [Weschera](https://www.contextstudios.ai/local-ai/creators/weschera.md)
- Engine: llama.cpp (prism fork)
- Quantization: Ternary PQ2_0 (empfohlen) / PTQ1_0 (5,9 GB); mmproj-Q8_0 für Vision
- Model family: Bonsai-2-27B (Qwen3.8-27B-Basis)
- Hardware: mac x1
- Context: 131072

## All numbers

- 10.6 tok/s (Everyday; Everyday, 1 request, prose prompt, no speculative decoding) — [source](https://github.com/Weschera/Bonsai-2-27B-Mac/blob/main/README.md)
- decode_tps: 10.6 (1 request) — [source](https://github.com/Weschera/Bonsai-2-27B-Mac/blob/main/README.md)
- decode_tps: 34.5 (1 request) — [source](https://github.com/Weschera/Bonsai-2-27B-Mac/blob/main/README.md)
- prefill_tps: 242 (1 request) — [source](https://github.com/Weschera/Bonsai-2-27B-Mac/blob/main/README.md)

## Weights

- [prism-ml/Ternary-Bonsai-2-27B-gguf](https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf) (main, apache-2.0)

[https://github.com/Weschera/Bonsai-2-27B-Mac](https://github.com/Weschera/Bonsai-2-27B-Mac)

## Related

- [Local AI on your hardware.](https://www.contextstudios.ai/local-ai.md)
