---
type: "Article"
title: "DeepSeek-V4-Flash-0731 304.6B FP8 with vLLM on 4× DGX Spark"
description: "DeepSeek-V4-Flash-0731 304.6B FP8 with vLLM on 4× DGX Spark: recipe by fujitsupolycom with speed, condition and source for every number, weights and license."
resource: "https://www.contextstudios.ai/local-ai/fujitsupolycom--deepseek-v4-flash-0731-tp4-4x-dgx-spark"
language: "en"
generated:
  by: "process:contextstudios-md/1"
  at: "2026-10-02T20:33:58.861Z"
status: "stable"
---

# DeepSeek-V4-Flash-0731 304.6B FP8 with vLLM on 4× DGX Spark

DeepSeek-V4-Flash-0731 304.6B FP8 with vLLM on 4× DGX Spark: 68.8 tok/s according to github.com (dataset as of Sep 29, 2026).

- Creator: [fujitsupolycom](https://www.contextstudios.ai/local-ai/creators/fujitsupolycom.md)
- Engine: vLLM
- Quantization: Stock-Checkpoint (FP8 block-quantisiert, fp8_ds_mla KV)
- Model family: DeepSeek-V4-Flash-0731
- Hardware: dgx-spark x4
- Context: 1048576
- Artificial Analysis Intelligence Index: 34.3 ([source](https://artificialanalysis.ai/models/deepseek-v4-flash))

## All numbers

- 68.8 tok/s (Everyday; Everyday, 1 request, realistic prompt, with DSpark (per recipe)) — [source](https://github.com/FujitsuPolycom/sparkring/blob/main/performance/records/deepseek-v4-flash/normalized-tp4-base-temp1-n5-20260823.md)
- decode_tps: 68.84 (1 request) — [source](https://github.com/FujitsuPolycom/sparkring/blob/main/performance/records/deepseek-v4-flash/normalized-tp4-base-temp1-n5-20260823.md)
- decode_tps: 105.88 (1 request) — [source](https://github.com/FujitsuPolycom/sparkring/blob/main/performance/records/deepseek-v4-flash/normalized-tp4-base-temp1-n5-20260823.md)
- prefill_tps: 2488.33 (1 request) — [source](https://github.com/FujitsuPolycom/sparkring/blob/main/performance/records/deepseek-v4-flash/normalized-tp4-base-temp1-n5-20260823.md)

## Weights

- [deepseek-ai/DeepSeek-V4-Flash-0731](https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731) (main, mit)

[https://github.com/FujitsuPolycom/sparkring](https://github.com/FujitsuPolycom/sparkring)

## Related

- [Local AI on your hardware.](https://www.contextstudios.ai/local-ai.md)
