1 request, mixed prompt set, with MTP (EAGLE/nextn) (per recipe) Source
Local AI · Recipe · 4× RTX 3090
Ornith-1.5-35B-A3B EXL3 3 bpw with ExLlamaV3 on 4× RTX 3090
by 0xSero
≤ 3 bit: quantization may cost quality
Engine
- Engine
- exllamav3 (TabbyAPI) 1.4.3 (+ Patch tp_import_split_n)
- Quantization
- EXL3 3 bpw Experten (MCG), Rest BF16, MTP 4 bpw
- Model family
- Ornith-1.5-35B-A3B
- Context
- 262,144
- Parameters
- 35B total / 3B active
- Creator
- 0xSero
- GitHub stars
- 0
- Repo updated
- Aug 26, 2026
Measurements
Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.
No sourced measurements for this recipe.
What you need
- Hardware
- 4 × NVIDIA RTX 3090
- Weights
- Main weights0xSero/Ornith-1.5-35B-A3B-EXL3-3bpw MIT
- Engine
- exllamav3 (TabbyAPI) 1.4.3 (+ Patch tp_import_split_n)
- Context
- 262,144 tokens
Notes
What matters before you rebuild it.
- ≤ 3 bit: quantization may cost qualityAt 3 bit and below the model may answer noticeably worse than the original. The intelligence number refers to the original.
Sources
Related recipes
1 request, prose prompt, with DFlash (per recipe) Source
- Experimental
1 request, prose prompt, with MTP (per recipe) Source
- License unclear
1 request, prose prompt, with MTP (per recipe) Source
1 request, prose prompt, with MTP (per recipe) Source
- Experimental
- License unclear
1 request, prose prompt, with DFlash2 (per recipe) Source
Local AI in your company?
In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.