1 request, synthetic prompt Source
Engine
- Engine
- llama.cpp
- Quantization
- GGUF Q4_K_M
- Model family
- gpt-oss
- Context
- —
- Parameters
- 21B MoE
- Creator
- notwitcheer
- GitHub stars
- 40
- Repo updated
- Sep 20, 2026
Measurements
Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.
No sourced measurements for this recipe.
What you need
- Hardware
- 1 × NVIDIA RTX 5090
- Weights
- Main weightsunsloth/gpt-oss-20b-GGUF Apache-2.0
- Engine
- llama.cpp
Related recipes
1× RTX PRO 6000
287tok/sPeak
Intelligence9with thinking
1× RTX 5090
283tok/sPeak
1 request, synthetic prompt Source
Intelligence9with thinking
1× RTX PRO 6000
196tok/sPeak
1 request, synthetic prompt Source
Intelligence12with thinking
1× RTX PRO 6000
171tok/sPeak
1 request, synthetic prompt Source
Intelligence12with thinking
1× RTX 3090
162tok/sPeak
1 request, synthetic prompt Source
Intelligence9with thinking
1× RTX 5090
46.5tok/sPeak
1 request, synthetic prompt Source
Intelligence12with thinking
- Partly in system RAM
Local AI in your company?
In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.