---
type: "GlossaryTerm"
title: "Layer Streaming"
description: "Layer Streaming is a memory technique for local inference of large language models. Model weights stay as a file on the NVMe SSD, and the engine loads only the"
resource: "https://www.contextstudios.ai/glossary/layer-streaming"
language: "en"
tags: ["infrastructure"]
generated:
  by: "process:contextstudios-md/1"
  at: "2026-10-08T20:45:16.324Z"
status: "stable"
---

# Layer Streaming

Layer Streaming is a memory technique for local inference of large language models. Model weights stay as a file on the NVMe SSD, and the engine loads only the layers needed for the next token into RAM, one forward pass at a time. This lets a 2.8 trillion parameter MoE model such as Kimi K3 (1.45 TB of weights) run on a MacBook with 128 GB RAM. The trade-off: sequential load overhead per token instead of a single full load. It pays off most for MoE models because only a subset of experts is active per token.
