Local AI · Recipe · 1× DGX Spark

Bonsai-27B Q4_K_M with llama.cpp on 1× DGX Spark

by vcruz305

15.7tok/sPeak

1 request, best value reported by source

Source: huggingface.co
Intelligence (original model) · Artificial Analysisno independent value

≤ 3 bit: quantization may cost quality

Engine

Engine
llama.cpp 62061f91 (prism) / mainline cecbf5fb0
Quantization
GGUF Q4_K_M (imatrix); Ladder Q3_K_M–Q8_0
Model family
Bonsai-27B
Context
16,384
Parameters
27B dense
Creator
vcruz305
GitHub stars
1
Repo updated
Jul 14, 2026

Measurements

Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.

No sourced measurements for this recipe.

What you need

Hardware
1 × NVIDIA DGX Spark (GB10)
Weights
Engine
llama.cpp 62061f91 (prism) / mainline cecbf5fb0
Context
16,384 tokens

Notes

What matters before you rebuild it.

  • ≤ 3 bit: quantization may cost qualityAt 3 bit and below the model may answer noticeably worse than the original. The intelligence number refers to the original.

Sources

Related recipes

← Back to overview

Local AI in your company?

In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.