LLMs
DeepSeek-V4-Flash Fits Because Its Experts Are 4-Bit
DeepSeek-V4-Flash activates 13B parameters per token but keeps 284B resident. Its experts ship in FP4, and that one choice sets almost the whole memory bill.
22 days ago
All articles about Mixture Of Experts