AI Infrastructure
Memory Bandwidth
- Definition
- Memory bandwidth is the amount of data an accelerator moves between its working memory and its compute cores per second, in gigabytes per second (GB/s). It is a throughput measure — distinct from capacity, which only counts bytes that fit. The chip generation (GDDR, HBM, LPDDR) sets the level: consumer cards deliver a few hundred GB/s, HBM accelerators one to two TB/s and beyond. For LLM inference, bandwidth is the single most important number, because decoding is bandwidth-bound: every generated token requires all weight matrices to stream through memory once. Tokens per second roughly equal bandwidth divided by the weight bytes. A 27-billion-parameter model at 8-bit occupies about 27 GB; at 178 GB/s that yields around 6.5 tokens per second — regardless of listed TFLOPS. Prefill and decode differ: prefill has high arithmetic intensity and benefits from raw compute; decode is decided by bandwidth. Local coding agents with long contexts show the pattern — the prompt vanishes quickly, the answer arrives at a low token rate. In practice: fix target token rate and model size, then pick the bandwidth class. Quantization (4-bit instead of 8-bit) halves weight bytes and directly raises the rate; batching spreads memory traffic over parallel tasks. If the model no longer fits, even high bandwidth stops helping — capacity becomes the limit. Memory bandwidth is the hinge between model architecture, hardware choice, and operating costs of local AI systems.
- Category
- AI Infrastructure
Deep Dive: Memory Bandwidth
Memory bandwidth is the amount of data an accelerator moves between its working memory and its compute cores per second, in gigabytes per second (GB/s). It is a throughput measure — distinct from capacity, which only counts bytes that fit. The chip generation (GDDR, HBM, LPDDR) sets the level: consumer cards deliver a few hundred GB/s, HBM accelerators one to two TB/s and beyond. For LLM inference, bandwidth is the single most important number, because decoding is bandwidth-bound: every generated token requires all weight matrices to stream through memory once. Tokens per second roughly equal bandwidth divided by the weight bytes. A 27-billion-parameter model at 8-bit occupies about 27 GB; at 178 GB/s that yields around 6.5 tokens per second — regardless of listed TFLOPS. Prefill and decode differ: prefill has high arithmetic intensity and benefits from raw compute; decode is decided by bandwidth. Local coding agents with long contexts show the pattern — the prompt vanishes quickly, the answer arrives at a low token rate. In practice: fix target token rate and model size, then pick the bandwidth class. Quantization (4-bit instead of 8-bit) halves weight bytes and directly raises the rate; batching spreads memory traffic over parallel tasks. If the model no longer fits, even high bandwidth stops helping — capacity becomes the limit. Memory bandwidth is the hinge between model architecture, hardware choice, and operating costs of local AI systems.
Production-Ready Guardrails