---
type: "Article"
title: "Qwen3.8-Flash-Next 125B with MLX on 1× Mac"
description: "Qwen3.8-Flash-Next 125B with MLX on 1× Mac: recipe by Weschera with speed, condition and source for every number, weights and license."
resource: "https://www.contextstudios.ai/local-ai/weschera--qwen38-flash-next-omlx-mtp6-mac-studio-m4-max"
language: "en"
generated:
  by: "process:contextstudios-md/1"
  at: "2026-10-02T12:55:06.229Z"
status: "stable"
---

# Qwen3.8-Flash-Next 125B with MLX on 1× Mac

Qwen3.8-Flash-Next 125B with MLX on 1× Mac: 83.1 tok/s according to github.com (dataset as of Sep 29, 2026).

- Creator: [Weschera](https://www.contextstudios.ai/local-ai/creators/weschera.md)
- Engine: oMLX 0.6.4 (custom kernels)
- Quantization: oQ4e gemischt (4-bit gesamt, group 64, MLX safetensors)
- Model family: Qwen3.8-Flash-Next
- Hardware: mac x1

## All numbers

- 83.1 tok/s (Peak; Peak, 1 request, code prompt) — [source](https://github.com/Weschera/qwen38-flash-next-omlx-mac/blob/main/README.md)
- decode_tps: 83.06 (1 request) — [source](https://github.com/Weschera/qwen38-flash-next-omlx-mac/blob/main/README.md)
- prefill_tps: 215 (1 request) — [source](https://github.com/Weschera/qwen38-flash-next-omlx-mac/blob/main/README.md)

## Weights

- [Jundot/Qwen3.8-Flash-Next-oQ4e-mtp](https://huggingface.co/Jundot/Qwen3.8-Flash-Next-oQ4e-mtp) (main)

[https://github.com/Weschera/qwen38-flash-next-omlx-mac](https://github.com/Weschera/qwen38-flash-next-omlx-mac)

## Related

- [Local AI on your hardware.](https://www.contextstudios.ai/local-ai.md)
