1 request, synthetic prompt Source
Local AI · Recipe · 1× RTX 5090
Llama 8B FP16 with vLLM on 1× RTX 5090
by DaytonLecoq
No speed value is documented for this recipe, so this page is not indexed.
Engine
- Engine
- vLLM nightly
- Quantization
- FP16
- Model family
- Llama
- Context
- —
- Parameters
- 8B dense
- Creator
- DaytonLecoq
- GitHub stars
- 0
- Repo updated
- May 11, 2026
Measurements
Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.
No sourced measurements for this recipe.
What you need
- Hardware
- 1 × NVIDIA RTX 5090
- Weights
- Main weightsmeta-llama/Llama-3.1-8B-Instruct llama3.1Access on request
- Engine
- vLLM nightly
Notes
What matters before you rebuild it.
- Smoke test
- Access on request
Sources
Related recipes
1 request, synthetic prompt Source
1 request, synthetic prompt Source
1 request, prose prompt, no speculative decoding Source
1 request, prose prompt, with MTP (per recipe) Source
1 request, mixed prompt set, with MTP Source
- Custom kernel required
Local AI in your company?
In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.