Local AI · Creator

jpezzulli

4 published recipes.

Recipes

1× RTX PRO 6000
89.8tok/sPeak

1 request, best value reported by source Source 

Intelligence34with thinkingOriginal model's value. Heavily compressed (~2 bit), intelligence is likely much lower.
  • ≤ 3 bit: quantization may cost quality

2-bit W2-Experten (MoET-Planes) + 6 GiB FP4-Korrektur-Tier (512×12 MiB), FP8-MLA-KV; Basis offizielles MXFP4-CheckpointvLLMWeights Repo Updated Aug 19, 2026

Details
1× RTX PRO 6000
85.6tok/sPeak

1 request, best value reported by source, context 500000 Source 

Intelligence34with thinkingOriginal model's value. Heavily compressed (~2 bit), intelligence is likely much lower.
  • ≤ 3 bit: quantization may cost quality

2-bit W2-Experten (MoET-Planes) + 6 GiB FP4-Korrektur-Tier (512×12 MiB), FP8-MLA-KV; Basis offizielles MXFP4-CheckpointvLLMWeights Repo Updated Aug 19, 2026

Details

← Back to overview