AI Infrastructure

Layer Streaming

Definition
Layer Streaming is a memory technique for local inference of large language models. Model weights stay as a file on the NVMe SSD, and the engine loads only the layers needed for the next token into RAM, one forward pass at a time. This lets a 2.8 trillion parameter MoE model such as Kimi K3 (1.45 TB of weights) run on a MacBook with 128 GB RAM. The trade-off: sequential load overhead per token instead of a single full load. It pays off most for MoE models because only a subset of experts is active per token.
Category
AI Infrastructure

Deep Dive: Layer Streaming

Layer Streaming is a memory technique for local inference of large language models. Model weights stay as a file on the NVMe SSD, and the engine loads only the layers needed for the next token into RAM, one forward pass at a time. This lets a 2.8 trillion parameter MoE model such as Kimi K3 (1.45 TB of weights) run on a MacBook with 128 GB RAM. The trade-off: sequential load overhead per token instead of a single full load. It pays off most for MoE models because only a subset of experts is active per token.

Tech Stack
DeltafinKimi K3

Production-Ready Guardrails