LLMs
Qwen 3.8 27B Hardware Guide: From the RTX 3090 to the DGX Spark
Qwen3.8-27B runs on a surprising range of hardware – but how fast, really? Baseline vs. community-tuned benchmarks (as of Aug 19, 2026): an RTX 3090 pushing ~1,000 tok/s across 64 parallel streams, an RTX 5090 at 148 tok/s, a DGX Spark at 210 tok/s aggregate, DFlash 2 at 3.4x on the H200 – plus RTX 5080 laptop, Mac mini, MacBook Pro, Mac Studio and AMD Strix Halo with prefill, decode and concurrency.
1 day ago