1 request, prose prompt, with MTP Source
- ≤ 3 bit: quantization may cost quality
Your goal
Be visible where AI answers
Automate processes
Build a product
Put AI agents to work
Connect and modernize systems
Know where we stand
Use Cases
CRMStrengthen customer relationshipsPopularE-CommerceBoost online revenueBooking System24/7 appointment bookingProject ManagementCoordinate teamsInvoicingGet paid fasterAnalyticsData-driven decisionsLocal AI · Recipe · 4× DGX Spark
GLM-5.2 753B INT4/INT8 with vLLM on 4× DGX Spark: 32.5 tok/s according to github.com (dataset as of Sep 29, 2026).
by Weschera
Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.
What matters before you rebuild it.
1 request, prose prompt, with MTP Source
1 request, prose prompt, with MTP Source
1 request, prose prompt, with MTP (per recipe) Source
1 request, realistic prompt, with MTP (fixed-4, greedy Draft) (per recipe) Source
1 request, mixed prompt set, with MTP (per recipe) Source
1 request, mixed prompt set, with MTP (per recipe) Source
In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.