---
type: "BlogPosting"
title: "Deltafin: 2.8T Kimi K3 from Four SSDs on a MacBook — Layer Streaming as a Local Inference Lever"
description: "The 2.8T MoE model Kimi K3 requires 1.45 TB of memory, exceeding any laptop's RAM. The Deltafin fork solves this by streaming expert weights from NVMe SSDs to RAM layer by layer. This article breaks down benchmarks on a 128GB M5 Max MacBook Pro, revealing true decode rates, SSD scaling, and practical limitations for local inference."
resource: "https://www.contextstudios.ai/blog/deltafin-2-8t-kimi-k3-from-four-ssds-on-a-macbook-layer-streaming"
language: "en"
tags: ["Local AI", "LLM", "Inference Optimization", "MoE", "Hardware"]
generated:
  by: "process:contextstudios-md/1"
  at: "2026-10-02T11:22:34.216Z"
status: "stable"
---

# Deltafin: 2.8T Kimi K3 from Four SSDs on a MacBook — Layer Streaming as a Local Inference Lever

Published: 2026-09-23
Tags: Local AI, LLM, Inference Optimization, MoE, Hardware

![Deltafin: 2.8T Kimi K3 from Four SSDs on a MacBook — Layer Streaming as a Local Inference Lever](https://wary-platypus-754.convex.cloud/api/storage/cae6b87f-4b18-4f39-8818-b9776e9b4463)

**TL;DR**
- The Mixture-of-Experts (MoE) model Kimi K3 comes with 1.45 TB of weights — no laptop RAM can hold that. The Deltafin fork offloads the experts to NVMe SSDs and streams them into RAM layer by layer.
- Benchmarks from September 8, 2026: 1.00 tokens per second decode on an M5 Max MacBook Pro with 128 GB and four SSDs — with a 6.3-minute wait to the first token after a 512-token prompt.
- The replicable core is a layer budget check in three calculation steps: It shows which model class runs on your own hardware, and how much additional SSDs actually accelerate the process.

## The Problem: Giant Models, Limited RAM

Moonshot's Kimi K3 is a Mixture-of-Experts (MoE) model with 2.8 trillion parameters. The expert weights comprise 1.45 TB — that's about eleven times more memory than a MacBook Pro with 128 GB can hold.

A standard loading process fails at this limit: The model doesn't fit entirely into memory. This is exactly where **layer streaming** comes in, as implemented by the [argonautlabsai/deltafin](https://github.com/argonautlabsai/deltafin) fork.

## The Solution: Layer-by-Layer Streaming from NVMe

The principle is simple: Hot layers remain in RAM, while the rest reside on up to four SSDs and are loaded on demand. Yet, every token is still decided by K3 itself — nothing is truncated, no expert is skipped.

All 16 experts per token are utilized, exactly as Moonshot shipped them. A small draft model is allowed to predict ahead, but K3 verifies each of these suggestions and confirms every token. The result is full model quality with a minimal RAM footprint.

The numbers speak an unusually honest language: Every value comes from a single cold run with an exact prompt; logs and placement manifests are completely open in the repo.

## The Benchmarks from September 8, 2026

| Metric | Drafter Off | Drafter On |
| --- | --- | --- |
| Decode rate, 512-token response (tok/s) | 0.92 | 1.00 |
| Decode rate, 128-token response (tok/s) | 0.93 | 1.13 |
| Public 17-token prompt, median of 3 (tok/s) | — | 0.96 |
| Time to first token, 512-token prompt | 6.3 min | 6.3 min |

For comparison: The upstream repo reports 0.2901 tok/s on an M1 Max — so the M5 Max combination with four SSDs is a good three times faster. The public 17-token prompt, at 0.96 tok/s, was clearly above the 0.68 reported upstream.

## SSD Scaling: More Drives, Less Linear

The four SSDs aren't a headline gimmick, but measurable physics. The scaling on identical prompts:

- **1 SSD:** ≈ 52% of the four-drive speed
- **2 SSDs:** ≈ 73%
- **3 SSDs:** ≈ 90%
- **4 SSDs:** 100% (Reference)

The reason: 16 reads run in parallel per layer, and the slowest set dictates the pace. Overall bandwidth only helps if the individual reads are fast enough.

## The Honest Limit: Prefill and Time-to-First-Token

The **decode rate** of ~1 tok/s is only half the story. A 512-token prompt takes 6.3 minutes to the first token (**Time-to-First-Token**, TTFT) — that is the true design limit.

The team identified the cause as the **prefill** pass re-reading the experts of each layer eight times. A fix is planned but wasn't built at the time of measurement. In practice, this means: short, precise prompts yield the best wait-time-to-response ratio.

## The Replicable Layer Budget Check

These three steps work with any MoE model and any hardware — without needing the repo:

1. **Note the total weight size.** Example K3: 1,450 GB. The number is in the Model Card or derived from the size of the weight files.
2. **Calculate the hot fraction.** Usable RAM divided by total weight: 128 / 1450 ≈ 9 percent of the weights remain permanently in RAM. This is the base from which streaming occurs.
3. **Apply the SSD scaling.** Read from 1.00 tok/s with four drives: one drive ≈ 0.52, two ≈ 0.73, three ≈ 0.90 tok/s. This allows you to roughly predict the performance of any custom configuration.

This check is especially worthwhile when comparing with a Cloud API: At ~1 tok/s, a 100-token response equals about 100 seconds of runtime — okay for asynchronous batch jobs, but not for interactive chat responses with multiple prompts per minute.

## What Does This Mean for the Agent Stack?

- **Large MoE flagships become locally executable** without having to squeeze 1.45 TB into RAM. Quality remains untouched with full expert utilization.
- **Latency is the new currency:** Anyone who doesn't factor in TTFT and decode rate will face a 6-minute traffic jam per prompt.
- **The budget check becomes the decision rule:** Only the 256 GB and multi-drive classes lift local inference out of the niche realm — exactly the configurations that are also the focus of the Apple hardware article on this blog.

## Conclusion

Deltafin proves that NVMe layer streaming genuinely bridges the gap between 128 GB of laptop RAM and a 1.45 TB model size — with clean, reproducible measurements instead of prose promises. The lever is the number of SSDs, the bottleneck is the prefill. For builders, this means: MoE flagships running locally are no longer just theory, but their speed must be calculated per workflow, not simply read off a chart.

## Frequently Asked Questions

**How many SSDs does Deltafin really need?**
The reference benchmark runs with four NVMe drives, achieving 1.00 tok/s there. A single drive still delivers about 52 percent of this speed, two drives 73 percent, and three 90 percent — so the jump from one to two is the most rewarding. If you only need one job per minute, two drives will serve you perfectly; for parallel responses, the full setup is worth it.

**Why is the first token so slow after 512-token prompts?**
The prefill pass reads the experts of each layer eight times, and each of these read series must finish before the first token is output. This adds up to a measured 6.3 minutes for a 512-token prompt. A fix has been announced in the repo but was not implemented at the time of measurement.

**Is 128 GB RAM worth it for local inference?**
With 128 GB, Kimi K3 can indeed be operated via layer streaming; the hot fraction of around 9 percent of the weights sustains the operation. The 256 GB class pushes the boundary further up and mitigates smaller prefill chains. The honest answer depends on the use case: For long prompts, the time to first token remains the bottleneck, not the RAM.

**How do I calculate the layer budget check for my own model?**
Divide your usable RAM by the total weight size from the model description, and you'll know your hot fraction. Then multiply the reference decode rate by the SSD scaling (0.52 / 0.73 / 0.90 / 1.00). The result is a rough but verifiable prediction — it can be cross-checked with the open logs in the repo's `k3-public-bench` folder.

**Does this replace a Cloud API?**
For asynchronous tasks at one token per second, this is highly usable locally, for instance, for note summarization or nightly batch runs. However, for interactive chains of five prompts with 300 tokens each, the runtime adds up to several minutes. The pragmatic mix remains: a large model locally as an anchor, and small, fast draft work in the cloud as usual.

## Sources

- [https://github.com/argonautlabsai/deltafin](https://github.com/argonautlabsai/deltafin) — Fork with benchmarks, measurement logs, and layer placement manifests
- [https://github.com/gavamedia/deltafin](https://github.com/gavamedia/deltafin) — Upstream engine (MIT), original by gavamedia
- [https://github.com/argonautlabsai/argodrive](https://github.com/argonautlabsai/argodrive) — Measurement and diagnostic tools (ARGODRIVE)
- [https://news.ycombinator.com/](https://news.ycombinator.com/) — Discussion of the run with 198 comments

## Related

- [LLM Development](https://www.contextstudios.ai/llm-development.md)
- [AI Development](https://www.contextstudios.ai/ai-development.md)
- [LLM Integration](https://www.contextstudios.ai/llm-integration.md)
- [AI Consulting](https://www.contextstudios.ai/ai-consulting.md)
