Local AI · Recipe · 1× DGX Spark

Laguna-S-2.1 118B NVFP4 with vLLM on 1× DGX Spark (128k context)

Laguna-S-2.1 118B NVFP4 with vLLM on 1× DGX Spark (128k context): 19.3 tok/s according to github.com (dataset as of Sep 29, 2026).

by Weschera

19.3tok/sEveryday

1 request, prose prompt, context 1k, no speculative decoding

Source: github.com
Intelligence (original model) · Artificial Analysisno independent value

Engine

Engine
vLLM
Quantization
NVFP4 (Blackwell-native, auto-detected aus quantization_config)
Model family
Laguna-S-2.1
Context
131,072
Parameters
118B total / 8B aktiv MoE (README)
Creator
Weschera
GitHub stars
3
Repo updated
Jul 21, 2026

Measurements

Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.

Decode · 1 request

  • Everyday (decode)19.3 tok/scontext 1K · Kein Spec-Decoding im Baseline-Profil; 19.0 @8K, 18.2 @31K — barely degrades Source 
  • Peak (decode)53 tok/sDFlash k=15, pure code generation (600-tok class impl), 2.8×; prose ~19 (kein Gain); mean acceptance ~2.3 auf Code Source 

Prefill by context

  • 31K3,703 tok/s2170 @1K, 3451 @8K, 3703 @31K Source 

What you need

Hardware
1 × NVIDIA DGX Spark (GB10)
Weights
Engine
vLLM
Context
131,072 tokens

Related recipes

← Back to overview

Local AI in your company?

In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.