Deltafin: 2.8T Kimi K3 from Four SSDs on a MacBook — Layer Streaming as a Local Inference Lever
(01)The 2.8T MoE model Kimi K3 requires 1.45 TB of memory, exceeding any laptop's RAM. The Deltafin fork solves this by streaming expert weights from NVMe SSDs to RAM layer by layer. This article breaks down benchmarks on a 128GB M5 Max MacBook Pro, revealing true decode rates, SSD scaling, and practical limitations for local inference.