Local AI · Recipe · 1× Mac

DeepSeek-V4-Flash 284B IQ2_XXS with DwarfStar on 1× Mac (vision)

DeepSeek-V4-Flash 284B IQ2_XXS with DwarfStar on 1× Mac (vision): 28.5 tok/s according to github.com (dataset as of Sep 29, 2026).

by Weschera

28.5tok/sEveryday

1 request, prose prompt, context 32k, with DSpark (per recipe)

Source: github.com
Intelligence (original model) · Artificial Analysisno independent value

≤ 3 bit: quantization may cost quality

Engine

Engine
ds4 (DwarfStar)
Quantization
IQ2_XXS imatrix q2 (antirez), w2Q2K/AProjQ8/SExpQ8/OutQ8; GGUF 86,720,111,776 bytes
Model family
DeepSeek-V4-Flash
Context
1,048,576
Parameters
284B / 13B aktiv + ViT (lt. Schwester-vLLM-Repo sfxnz/Vision-Exp; hier nicht separat belegt)
Creator
Weschera
GitHub stars
2
Repo updated
Sep 2, 2026

Measurements

Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.

Decode · 1 request

  • Everyday (decode)28.5 tok/scontext 32K · Thinking off, 404 tokens Source 
  • Peak (decode)31.7 tok/scontext 32K · Thinking off, code prompt, 1,656 tokens; @1M ctx: 32.0 Source 

What you need

Hardware
1 × Mac / Apple Silicon
Engine
ds4 (DwarfStar)
Context
1,048,576 tokens

Notes

What matters before you rebuild it.

  • ≤ 3 bit: quantization may cost qualityAt 3 bit and below the model may answer noticeably worse than the original. The intelligence number refers to the original.

Sources

Related recipes

← Back to overview

Local AI in your company?

In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.