LLMs
DeepSeek-V4-Flash Fits Because Its Experts Are 4-Bit
DeepSeek-V4-Flash activates 13B parameters per token but keeps 284B resident. Its experts ship in FP4, and that one choice sets almost the whole memory bill.
19 days ago
All articles about Llm Ops
DeepSeek-V4-Flash activates 13B parameters per token but keeps 284B resident. Its experts ship in FP4, and that one choice sets almost the whole memory bill.
Two API settings moved GPT-5.6 Sol from 13.3% to 38.3% on ARC-AGI-3 with no weight change. Why the harness, not the model, decided that benchmark score.
GPT-5.6 Pro is unannounced and every spec is a leak. A builder's checklist to make any frontier launch a non-event: pin model IDs, freeze evals, cap cost, name a fallback.