---
type: "Article"
title: "Qwen3.8-Flash-Next 125B mit MLX auf 1× Mac"
description: "Qwen3.8-Flash-Next 125B mit MLX auf 1× Mac: Rezept von Weschera mit Tempo, Bedingung und Quelle je Zahl, Gewichten und Lizenz."
resource: "https://www.contextstudios.ai/de/lokale-ki/weschera--qwen38-flash-next-omlx-mtp6-mac-studio-m4-max"
language: "de"
generated:
  by: "process:contextstudios-md/1"
  at: "2026-10-02T23:08:36.987Z"
status: "stable"
---

# Qwen3.8-Flash-Next 125B mit MLX auf 1× Mac

Qwen3.8-Flash-Next 125B mit MLX auf 1× Mac: 83,1 tok/s laut github.com (Datensatz-Stand: 29. Sept. 2026).

- Kreator: [Weschera](https://www.contextstudios.ai/de/lokale-ki/kreatoren/weschera.md)
- Engine: oMLX 0.6.4 (custom kernels)
- Quantisierung: oQ4e gemischt (4-bit gesamt, group 64, MLX safetensors)
- Modellfamilie: Qwen3.8-Flash-Next
- Hardware: mac x1

## Alle Zahlen

- 83,1 tok/s (Spitze; Spitze, 1 Anfrage, Code-Prompt) — [source](https://github.com/Weschera/qwen38-flash-next-omlx-mac/blob/main/README.md)
- decode_tps: 83.06 (1 Anfrage) — [source](https://github.com/Weschera/qwen38-flash-next-omlx-mac/blob/main/README.md)
- prefill_tps: 215 (1 Anfrage) — [source](https://github.com/Weschera/qwen38-flash-next-omlx-mac/blob/main/README.md)

## Gewichte

- [Jundot/Qwen3.8-Flash-Next-oQ4e-mtp](https://huggingface.co/Jundot/Qwen3.8-Flash-Next-oQ4e-mtp) (main)

[https://github.com/Weschera/qwen38-flash-next-omlx-mac](https://github.com/Weschera/qwen38-flash-next-omlx-mac)

## Related

- [Lokale KI auf Ihrer Hardware.](https://www.contextstudios.ai/de/lokale-ki.md)
