---
type: "BlogPosting"
title: "TensorFold: The engine that makes local LLMs 2–3× faster — without a single token changing"
description: "Speculative decoding with a hard contract: decode faster without a single byte changing. What TensorFold can do, what it costs, and who should use it now."
resource: "https://www.contextstudios.ai/blog/tensorfold-the-engine-that-makes-local-llms-2-3-faster-without-a-single"
language: "en"
tags: ["Local AI", "Inference", "TensorFold", "DGX Spark"]
generated:
  by: "process:contextstudios-md/1"
  at: "2026-10-02T08:19:08.540Z"
status: "stable"
---

# TensorFold: The engine that makes local LLMs 2–3× faster — without a single token changing

Published: 2026-10-02
Tags: Local AI, Inference, TensorFold, DGX Spark

![TensorFold: The engine that makes local LLMs 2–3× faster — without a single token changing](https://wary-platypus-754.convex.cloud/api/storage/d01fdfae-cde4-47b0-9033-9bb1055e5c1a)

On Wednesday evening, Ash Hart posts on X: "I didn't imagine it to blow up like this." Three days earlier, his project TensorFold had crossed the 800-star mark on GitHub. That same day, version 0.6.1 ships — the twelfth release in five days. What began as a weekend experiment "to make Qwen 27B faster" has become the fastest-growing inference engine in the local-AI scene. The reason is a promise nobody else makes in this form: **TensorFold is faster, without a single byte of the output changing.**

## What is TensorFold?

TensorFold is an open-source inference engine (Apache-2.0) that serves local LLMs behind an OpenAI-compatible endpoint — on Apple Silicon (Metal/MLX) just as on NVIDIA GPUs including DGX Spark. One command is all it takes:

```bash
pip install git+https://github.com/ashhart/TensorFold.git
tensorfold serve Vontra/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-MLX-4bit --context 65536
```

The server downloads the model from Hugging Face, builds family-specific Metal or CUDA kernels, and exposes `/v1/chat/completions`. Any OpenAI client — coding agents, SDKs, curl — works without modification.

Behind the project is Ash Hart, an independent AI systems architect from the UK who has also built MCDMA (Metal-CUDA Direct Memory Access) and Imprint. He works on the project alongside a full-time job, now supported by two well-known names in the scene: **MiaAI-Lab** and **Volatile Markets** have officially joined the core team.

## The differentiator: byte-exact drafted decoding

The speed comes from speculative decoding: a small draft model (DFlash2 or an MTP head) guesses several tokens ahead, and the large model verifies them in batched "lanes" — multiple draft rows in a single forward pass. That alone isn't new; vLLM can do MTP too.

What's new is the accuracy contract. TensorFold guarantees that a drafted reply is **byte-identical** to the same request without drafts. The project writes its own row-exact kernels for this: the bits of a row may not depend on how many rows share the verify pass. Sampling is keyed — the token at position p is derived deterministically from seed, position and token id. The formula in the recipe book: *"Drafted decoding writes the same bytes as serial decoding. Drafts change speed only."*

Every release checks this empirically: every drafted reply is compared byte for byte against the same request with `"draft": false`. The server reports a `token_sha` hash and `min_rows` per reply. If you want to test it yourself, send the same prompt twice — once with, once without drafts — and diff the outputs.

## The numbers

On DGX Spark (GB10), TensorFold lands 1.6 to 3 times above vLLM with MTP depending on the model — measured through the same OpenAI client, median over five seeds:

| Model | Sparks | vs. vLLM (MTP=3) |
|---|---|---|
| Qwen3.8-27B + DFlash2 | 1 | 2.70–3.05× |
| Qwen3.8-27B + DFlash2 | 2 | 1.94–2.49× |
| Qwen3.8 Flash Next | 1 | 1.60–1.79× |
| Qwen3.8 Flash Next | 2 | 1.74–2.24× |
| GLM-5.3-Flash | 2 | 1.78–2.06× |

Concretely: GLM-5.3-Flash on two Sparks decodes at 63 tok/s; Qwen3.8-27B on a single Spark manages 58.3 tok/s of prose single-stream and 220 tok/s at 8 streams. On an M3 Ultra (without the M5's tensor units), Qwen3.8-27B went from 27 to as much as 131–160 tok/s.

Important context: this is not a quantization trick. Same 4-bit checkpoints, same quality — the engine simply extracts more from every weight read.

## The community does the rest

What sets TensorFold apart from comparable projects is the pace and the openness. In the last few days alone: RTX 40 support (Ada), RTX PRO 6000 Blackwell with NVFP4 checkpoints, the Responses API, Prometheus metrics, the switch to Apache-2.0 — and a stream of community PRs that get individually thanked by name in the release notes.

MiaAI-Lab has published a whole recipe collection on top of TensorFold: GLM-5.3-Flash EXL3 on 2× DGX Spark (with 1M context, 4 streams, a 2.9M-token KV pool), Qwen3.8-27B on a single Spark, and by its own announcement soon 3× and 4× Spark recipes for DeepSeek v4.1 Flash. Every recipe ships with a dedicated web page and measured numbers — and we've already added the GLM-5.3-EXL3 recipe to our [Local-AI recipe database](/local-ai).

Our own stress test: a 2× GB10 cluster (ASUS GX10) over a MikroTik CRS504 switch instead of a direct cable, GPUs capped at 2,200 MHz, reproducing the README numbers with sparkDash. The result: decode within ±5% of the reference, structured output at 2 and 4 streams even +11–14% above it, prefill within 3–5%, full 1M context. The recipe numbers reproduce outside the original hardware — a switch instead of a cable costs measurably almost nothing.

## The honest limits

Three things worth knowing:

**Concurrency is the trade-off.** A community benchmark from WescheNex1q puts it soberly: at 16 parallel requests on a Spark, vLLM clearly beats TensorFold (253 vs. 62 tok/s aggregate); at a single request, TensorFold wins (83 vs. 57). TensorFold is the fastest single-user engine there is — for team servers with dozens of concurrent agents, vLLM remains the choice. The current releases work on exactly this (prompts that fill inside the decode rounds, background priorities), but it's an open battle.

**Prefill doesn't get faster on Macs.** The speed gain lives in decode; on Apple Silicon, prefill throughput is comparable to other engines. On CUDA it looks much better thanks to CUDA graphs and kernel optimizations (up to 1.27× vLLM on long prompts).

**One maintainer with a day job.** Seven releases in 34 hours over the launch weekend — impressive, but also a bus factor of 1. Bringing in MiaAI-Lab and Volatile Markets is the direct answer to that, and community PRs keep coming. For production use it still means: pin versions, read changelogs.

Another often-overlooked point: TensorFold measures its own quality. For NVFP4 checkpoints the project publishes the top-1 agreement against an fp32 reference forward (92.9% in checkpoint math vs. vLLM's 93.0%) — instead of just celebrating tok/s. That fits a project whose whole promise is "same bits".

## Who is TensorFold for?

- **DGX Spark and Mac owners**: a clear win, with the same server and same output on both platforms. On a 2-Spark setup, GLM-5.3-Flash runs with 1M context on consumer hardware.
- **Coding-agent users**: streamed tool-call arguments, the Responses API, prompt caching (resent 18k prompts: from seconds to under 0.1 s) — exactly what agent loops need.
- **Team-server operators**: wait a while longer or run vLLM. The concurrency work is underway but not finished.

TensorFold is showing how fast local inference can move when someone writes the kernels themselves instead of configuring framework defaults. In 2025, the local-AI scene learned to run models locally. In 2026 it's learning to run them fast — and TensorFold is currently the most prominent example. Which checkpoints and recipes exist for which hardware is what our [Local-AI hub](/local-ai) surfaces with its public recipe database — including tested measurements instead of marketing numbers.

## FAQ

**Is TensorFold a model or an engine?**
An engine. TensorFold replaces the inference server (vLLM, SGLang or ollama, for example) and loads existing checkpoints from Hugging Face. The model itself — Qwen3.8-27B, GLM-5.3-Flash, Nemotron 3.5 Lightning, Qwen3.8 Flash Next — stays unchanged; the weights, and with them the intelligence, are exactly the same as on any other server. Only the way the kernels compute them is different.

**How can accelerated decoding be byte-identical?**
The trick is in the verification: the draft model only guesses, the large model checks and keeps only the guessed tokens it would have chosen itself. TensorFold writes kernels where a row in a batch yields the same bits as a single-step decode, and makes sampling deterministic from seed, position and token id. The result is automatically diffed against draft-free decoding in every release; the server additionally reports a token hash per reply.

**Do I need a DGX Spark?**
No. TensorFold runs on Apple Silicon (M1 through M5, each generation covered by its own kernels) and on NVIDIA GPUs from the RTX 40 series over DGX Spark (GB10) to the RTX PRO 6000 Blackwell. On a Mac, a pip install is enough; on Linux it runs cleanest in NVIDIA's PyTorch container. The most spectacular recipe numbers come from the Spark community, but a single RTX 4090 can already run Qwen3.8-27B with DFlash2 in a 40k context window.

**What about quality with the quantization?**
The checkpoints TensorFold uses (MLX 4-bit, EXL3/TR3 4bpw, NVFP4) are created by third parties and are independent of the engine. TensorFold doesn't change them — the same checkpoint file delivers the same quality as under vLLM. The project also measures itself: with NVFP4, the top-1 agreement against an fp32 reference run is 92.9% (vLLM: 93.0%), and 98.9% in full-precision mode. If you want maximum quality, run `--precision full` against the same weights.

**Can I run TensorFold for multiple users at once?**
To a limited degree. The focus is on a few low-latency streams — 4 to 8 parallel sessions are the sweet spot on DGX Spark, and the recipes are tuned for that. With dozens of simultaneous requests, the aggregate currently loses to vLLM, because vLLM's continuous batching is optimized for high density. The current releases ("prompts fill inside the decode rounds", background priorities) are working on exactly this bottleneck, and the gap narrows from release to release.

**How do I install it on a DGX Spark?**
The most comfortable way is a ready-made community recipe: MiaAI-Lab maintains a Docker setup (ghcr.io) that configures TensorFold with tensor parallelism across two Sparks, RoCE networking and the KV pool — you only put your SSH credentials into a local.sh and run start.sh. If you prefer manual: start NVIDIA's PyTorch container, pip-install TensorFold, pull the checkpoints with tensorfold pull, then start rank 1 (first) and rank 0 on both Sparks. The AI-agent runbooks in the repo describe every step so that a coding agent can carry out the setup on its own.

## Sources

- Ash Hart, TensorFold — GitHub repo: https://github.com/ashhart/TensorFold
- TensorFold release notes v0.4.0 to v0.6.1 (Sep 27 – Oct 1, 2026): https://github.com/ashhart/TensorFold/releases
- The CUDA recipe book — numbers against vLLM: https://github.com/ashhart/TensorFold/blob/main/docs/recipes/cuda.md
- GLM-5.3-Flash recipe (2× DGX Spark): https://github.com/ashhart/TensorFold/blob/main/docs/recipes/glm-5.3-flash.md
- MiaAI-Lab, GLM-5.3-Flash-EXL3-2x-DGX-Sparks-TensorFold: https://github.com/MiaAI-Lab/GLM-5.3-Flash-EXL3-2x-DGX-Sparks-TensorFold
- Ash Hart on X (@ashxhart), 0.6.0/0.6.1 announcements and team news (Sep 30 – Oct 1, 2026): https://x.com/ashxhart
- MiaAI-Lab on X (@MiaAI_lab), recipe announcements and KV-pool update (Oct 1, 2026): https://x.com/MiaAI_lab
- WescheNex1q, concurrency benchmark vLLM vs. SGLang vs. TensorFold (Oct 1, 2026): https://x.com/WescheNex1q
- SingularityByte, "TensorFold vs MTPLX vs mlx-dspark": https://singularitybyte.com/tools/exact-speculative-decoding-apple-silicon.html
- Market Intelligence Research, "Faster, provably the same model — TensorFold exact decoding": https://marketintelligenceresearch.com/blog/faster-provably-same-model-tensorfold-exact-decoding/
- LLMKube, "A weekend with TensorFold on a MacBook": https://llmkube.com/blog/tensorfold-m5-max-engine-not-quant
- Own measurement, 2× GB10 (ASUS GX10) via CRS504 switch, sparkDash 1.8.9, Oct 1, 2026 (reproducing the MiaAI recipe v1.3.1 on TensorFold v0.6.0)
- TensorFold — official site: https://tensorfold.dev/

## Related

- [AI Consulting](https://www.contextstudios.ai/ai-consulting.md)
- [AI Agency](https://www.contextstudios.ai/ai-development-company.md)
- [AI Development](https://www.contextstudios.ai/ai-development.md)
