LLMs
DeepSeek-V4-Flash Fits Because Its Experts Are 4-Bit
DeepSeek-V4-Flash activates 13B parameters per token but keeps 284B resident. Its experts ship in FP4, and that one choice sets almost the whole memory bill.
9 days ago
All articles about Quantization