AI Infrastructure
Layer Streaming
- Definition
- Layer Streaming is a memory technique for local inference of large language models. Model weights stay as a file on the NVMe SSD, and the engine loads only the layers needed for the next token into RAM, one forward pass at a time. This lets a 2.8 trillion parameter MoE model such as Kimi K3 (1.45 TB of weights) run on a MacBook with 128 GB RAM. The trade-off: sequential load overhead per token instead of a single full load. It pays off most for MoE models because only a subset of experts is active per token.
- Category
- AI Infrastructure
Deep Dive: Layer Streaming
Layer Streaming is a memory technique for local inference of large language models. Model weights stay as a file on the NVMe SSD, and the engine loads only the layers needed for the next token into RAM, one forward pass at a time. This lets a 2.8 trillion parameter MoE model such as Kimi K3 (1.45 TB of weights) run on a MacBook with 128 GB RAM. The trade-off: sequential load overhead per token instead of a single full load. It pays off most for MoE models because only a subset of experts is active per token.
Production-Ready Guardrails