Local AI · Recipe · 1× Mac

DeepSeek-V4.1-Flash 763B with MLX on 1× Mac

by drowzeys

25.1tok/sEveryday

1 request, prose prompt, short prompt, with MTP (DSpark built-in) (per recipe)

Source: github.com
Intelligence (original model) · Artificial Analysis40with thinkingno thinking 25

≤ 3 bit: quantization may cost quality

Engine

Engine
oMLX 0.7.0.dev2 @395ec2fd, MLX 0.32.2
Quantization
oQ3e (affine 3-bit) mit 27/40 MoE-Layern 2-bit gs64
Model family
DeepSeek-V4.1-Flash
Context
404,805
Parameters
763B MoE
Creator
drowzeys
GitHub stars
6
Repo updated
Sep 16, 2026

Measurements

Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.

No sourced measurements for this recipe.

What you need

Hardware
1 × Mac / Apple Silicon
Engine
oMLX 0.7.0.dev2 @395ec2fd, MLX 0.32.2
Context
404,805 tokens

Notes

What matters before you rebuild it.

  • ≤ 3 bit: quantization may cost qualityAt 3 bit and below the model may answer noticeably worse than the original. The intelligence number refers to the original.
  • No license on HF
  • License unclear

Sources

Related recipes

← Back to overview

Local AI in your company?

In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.