1 request, prose prompt, with DFlash2 (per recipe) Source
- Custom kernel required
- Non-commercial
Your goal
Be visible where AI answers
Automate processes
Build a product
Put AI agents to work
Connect and modernize systems
Know where we stand
Use Cases
CRMStrengthen customer relationshipsPopularE-CommerceBoost online revenueBooking System24/7 appointment bookingProject ManagementCoordinate teamsInvoicingGet paid fasterAnalyticsData-driven decisionsLocal AI · Recipe · 2× DGX Spark
GLM-5.3-Flash NVFP4 with vLLM on 2× DGX Spark (128k context, with DFlash2): 20.5 tok/s according to github.com (dataset as of Sep 29, 2026).
by Weschera
Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.
1 request, prose prompt, with DFlash2 (per recipe) Source
1 request, prose prompt, with DFlash2 (per recipe) Source
1 request, prose prompt, with DFlash2 (per recipe) Source
HumanEval 97.6% / GSM8K 98.0% with FP8 KV and 4-bit dense Source
1 request, prose prompt, with DFlash2 (per recipe) Source
1 request, prose prompt, with MTP4 Source
In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.