---
type: "Article"
title: "Bonsai-2-27B Q8_0 mit llama.cpp auf 1× Mac"
description: "Bonsai-2-27B Q8_0 mit llama.cpp auf 1× Mac: Rezept von Weschera mit Tempo, Bedingung und Quelle je Zahl, Gewichten und Lizenz."
resource: "https://www.contextstudios.ai/de/lokale-ki/weschera--bonsai-2-27b-pq2_0-prism-llamacpp-mac-studio-m4-max"
language: "de"
generated:
  by: "process:contextstudios-md/1"
  at: "2026-10-02T23:20:11.406Z"
status: "stable"
---

# Bonsai-2-27B Q8_0 mit llama.cpp auf 1× Mac

Bonsai-2-27B Q8_0 mit llama.cpp auf 1× Mac: 10,6 tok/s laut github.com (Datensatz-Stand: 29. Sept. 2026).

- Kreator: [Weschera](https://www.contextstudios.ai/de/lokale-ki/kreatoren/weschera.md)
- Engine: llama.cpp (prism fork)
- Quantisierung: Ternary PQ2_0 (empfohlen) / PTQ1_0 (5,9 GB); mmproj-Q8_0 für Vision
- Modellfamilie: Bonsai-2-27B (Qwen3.8-27B-Basis)
- Hardware: mac x1
- Kontext: 131072

## Alle Zahlen

- 10,6 tok/s (Alltag; Alltag, 1 Anfrage, Prosa-Prompt, ohne Spekulation) — [source](https://github.com/Weschera/Bonsai-2-27B-Mac/blob/main/README.md)
- decode_tps: 10.6 (1 Anfrage) — [source](https://github.com/Weschera/Bonsai-2-27B-Mac/blob/main/README.md)
- decode_tps: 34.5 (1 Anfrage) — [source](https://github.com/Weschera/Bonsai-2-27B-Mac/blob/main/README.md)
- prefill_tps: 242 (1 Anfrage) — [source](https://github.com/Weschera/Bonsai-2-27B-Mac/blob/main/README.md)

## Gewichte

- [prism-ml/Ternary-Bonsai-2-27B-gguf](https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf) (main, apache-2.0)

[https://github.com/Weschera/Bonsai-2-27B-Mac](https://github.com/Weschera/Bonsai-2-27B-Mac)

## Related

- [Lokale KI auf Ihrer Hardware.](https://www.contextstudios.ai/de/lokale-ki.md)
