---
type: "ProfilePage"
title: "Weschera"
description: "24 recipes by Weschera for local models, with speed, source and weights."
resource: "https://www.contextstudios.ai/local-ai/creators/weschera"
language: "en"
generated:
  by: "process:contextstudios-md/1"
  at: "2026-10-02T20:31:18.761Z"
status: "stable"
---

# Weschera


- [Qwen3.8-27B with MLX on 1× Mac](https://www.contextstudios.ai/local-ai/weschera--qwen38-27b-omlx-ane-mtp3-mac-studio-m4-max.md): 53.3 tok/s (Everyday; Everyday, 1 request, prose prompt, with MTP (nativ) (per recipe)) — [source](https://github.com/Weschera/Qwen3.8-27B-oMLX-MTP-Mac/blob/main/README.md)
- [Qwen3.8-27B NVFP4 with vLLM on 1× DGX Spark (Weschera)](https://www.contextstudios.ai/local-ai/weschera--qwen38-27b-nvfp4-unsloth-mtp3-1x-dgx-spark.md): 45 tok/s (Everyday; Everyday, 1 request, prose prompt, with MTP) — [source](https://github.com/Weschera/Qwen3.8-27B-DGX-Spark-Quant-Ladder/blob/main/README.md)
- [Qwen3.8-Flash-Next NVFP4 with vLLM on 2× DGX Spark (Weschera)](https://www.contextstudios.ai/local-ai/weschera--qwen38-flash-next-nvfp4-nvidia-mtp3-2x-dgx-spark.md): 33 tok/s (Everyday; Everyday, 1 request, prose prompt, with MTP) — [source](https://github.com/Weschera/Qwen3.8-Flash-Next-NVFP4-2x-DGX-Spark/blob/main/README.md)
- [DeepSeek-V4-Flash 284B IQ2_XXS with DwarfStar on 1× Mac (vision)](https://www.contextstudios.ai/local-ai/weschera--dsv4-flash-vision-exp-iq2xxs-dspark-ds4-mac-studio-m4-max.md): 28.5 tok/s (Everyday; Everyday, 1 request, prose prompt, context 32k, with DSpark (per recipe)) — [source](https://github.com/Weschera/DeepSeek-V4-Flash-Vision-Exp-ds4-Mac-Studio/blob/main/README.md)
- [Ling-3.0-flash-VL FP8 with SGLang on 2× DGX Spark](https://www.contextstudios.ai/local-ai/weschera--ling30-flash-vl-fp8-sglang-tp2-2x-dgx-spark.md): 26 tok/s (Everyday; Everyday, 1 request, prose prompt, no speculative decoding) — [source](https://github.com/Weschera/Ling-3.0-flash-VL-2x-DGX-Sparks/blob/main/README.md)
- [Qwen3.8-Flash-Next 125B UD-Q4_K_XL with llama.cpp on 1× DGX Spark](https://www.contextstudios.ai/local-ai/weschera--qwen38-flash-next-ud-q4kxl-mtp4-1x-dgx-spark.md): 24.2 tok/s (Everyday; Everyday, 1 request, prose prompt, with Draft) — [source](https://github.com/Weschera/Qwen3.8-Flash-Next-1x-DGX-Spark/blob/main/README.md)
- [GLM-5.3-Flash NVFP4 with vLLM on 2× DGX Spark (128k context, with MTP)](https://www.contextstudios.ai/local-ai/weschera--glm53-flash-nvfp4-vcruz-mtp4-2x-dgx-spark.md): 23 tok/s (Everyday; Everyday, 1 request, prose prompt, with MTP4) — [source](https://github.com/Weschera/glm53-flash-2spark-tp2/blob/main/README.md)
- [Spark-X2.5-4B BF16 with SGLang on 1× DGX Spark](https://www.contextstudios.ai/local-ai/weschera--spark-x25-4b-bf16-sglang-1x-dgx-spark.md): 21.8 tok/s (Everyday; Everyday, 1 request, prose prompt, no speculative decoding) — [source](https://github.com/Weschera/spark-x25-4b-dgx-recipe/blob/main/README.md)
- [GLM-5.3-Flash NVFP4 with vLLM on 2× DGX Spark (128k context, with DFlash2)](https://www.contextstudios.ai/local-ai/weschera--glm53-flash-nvfp4-nvidia-dflash2-k7-2x-dgx-spark.md): 20.5 tok/s (Everyday; Everyday, 1 request, prose prompt, no speculative decoding) — [source](https://github.com/Weschera/GLM-5.3-Flash-NVIDIA-NVFP4-2x-DGX-Spark/blob/main/README.md)
- [Laguna-S-2.1 118B NVFP4 with vLLM on 1× DGX Spark (128k context)](https://www.contextstudios.ai/local-ai/weschera--laguna-s-21-nvfp4-dflash15-1x-dgx-spark.md): 19.3 tok/s (Everyday; Everyday, 1 request, prose prompt, context 1k, no speculative decoding) — [source](https://github.com/Weschera/Laguna-S-2.1-NVFP4-1x-DGX-Spark/blob/main/README.md)
- [Qwen3.8-27B NVFP4 with Atlas on 1× DGX Spark](https://www.contextstudios.ai/local-ai/weschera--qwen38-27b-nvfp4-atlas-mtp3-1x-dgx-spark.md): 19 tok/s (Everyday; Everyday, 1 request, prose prompt, with MTP (per recipe)) — [source](https://github.com/Weschera/Atlas-DGX-Spark-Quickstart/blob/main/README.md)
- [GLM-5.3-Flash 320B UD-Q2_K_XL with llama.cpp on 1× DGX Spark](https://www.contextstudios.ai/local-ai/weschera--glm53-flash-ud-q2kxl-mtp-1x-dgx-spark.md): 17.9 tok/s (Everyday; Everyday, 1 request, prose prompt, context 32k, no speculative decoding) — [source](https://github.com/Weschera/glm53-flash-one-spark/blob/main/README.md)
- [Bonsai-2-27B Q8_0 with llama.cpp on 1× Mac](https://www.contextstudios.ai/local-ai/weschera--bonsai-2-27b-pq2_0-prism-llamacpp-mac-studio-m4-max.md): 10.6 tok/s (Everyday; Everyday, 1 request, prose prompt, no speculative decoding) — [source](https://github.com/Weschera/Bonsai-2-27B-Mac/blob/main/README.md)
- [DeepSeek-V4-Flash FP8 with vLLM on 2× DGX Spark (1024k context)](https://www.contextstudios.ai/local-ai/weschera--dsv4-flash-0731-dspark-k7-2x-dgx-spark.md): 83.8 tok/s (Mixed; Mixed, 1 request, mixed prompt set, with Draft) — [source](https://github.com/Weschera/DeepSeek-V4-Flash-0731-DSpark-2x-DGX-Spark/blob/main/README.md)
- [Qwen3.8-Flash-Next NVFP4 with SGLang on 2× DGX Spark](https://www.contextstudios.ai/local-ai/weschera--qwen38-flash-next-nvfp4-radixark-nextn-2x-dgx-spark.md): 41.7 tok/s (Mixed; Mixed, 1 request, mixed prompt set, with NEXTN (MTP) (per recipe)) — [source](https://github.com/Weschera/qwen38-flashnext-dgx-spark/blob/main/README.md)
- [Qwen3.8-Flash-Next 80B UD-IQ4_XS with llama.cpp on 1× DGX Spark](https://www.contextstudios.ai/local-ai/weschera--qwen38-flash-next-ud-iq4xs-262k-1x-dgx-spark.md): 26.9 tok/s (Mixed; Mixed, 1 request, mixed prompt set, no speculative decoding) — [source](https://github.com/Weschera/qwen38-flashnext-single-spark/blob/main/README.md)
- [Qwen3.8-27B NVFP4 with SGLang on 1× DGX Spark (Weschera)](https://www.contextstudios.ai/local-ai/weschera--qwen38-27b-nvfp4-dflash2-1x-dgx-spark.md): 25.2 tok/s (Mixed; Mixed, 1 request, mixed prompt set, with DFlash2 (per recipe)) — [source](https://github.com/Weschera/Qwen3.8-27B-NVFP4-DFlash2-DGX-Spark/blob/main/README.md)
- [GLM-5.3-Flash 320B UD-IQ3_XXS with llama.cpp on 1× DGX Spark](https://www.contextstudios.ai/local-ai/weschera--glm53-flash-ud-iq3xxs-mtp2-1x-dgx-spark.md): 20.8 tok/s (Mixed; Mixed, 1 request, mixed prompt set, with MTP) — [source](https://github.com/Weschera/GLM-5.3-Flash-Unsloth-1x-DGX-Spark/blob/main/README.md)
- [Qwen3.6-35B NVFP4 with SGLang on 1× DGX Spark](https://www.contextstudios.ai/local-ai/weschera--qwen36-35b-a3b-nvfp4-nextn-sglang-1x-dgx-spark.md): 92.2 tok/s (Peak; Peak, 1 request, best value reported by source, context 64k) — [source](https://github.com/Weschera/qwen-sglang-dgx-spark/blob/main/README.md)
- [Qwen3.8-Flash-Next 125B with MLX on 1× Mac](https://www.contextstudios.ai/local-ai/weschera--qwen38-flash-next-omlx-mtp6-mac-studio-m4-max.md): 83.1 tok/s (Peak; Peak, 1 request, code prompt) — [source](https://github.com/Weschera/qwen38-flash-next-omlx-mac/blob/main/README.md)
- [DeepSeek-V4-Flash NVFP4 with vLLM on 2× DGX Spark](https://www.contextstudios.ai/local-ai/weschera--dsv4-flash-dspark-1m-nvfp4kv-2x-dgx-spark.md): 65.4 tok/s (Peak; Peak, 1 request, best value reported by source) — [source](https://github.com/Weschera/DeepSeek-v4-Flash-DSpark-1M-NVFP4-KV-2x-DGX-Spark/blob/main/README.md)
- [GLM-5.2 753B INT4/INT8 with vLLM on 4× DGX Spark](https://www.contextstudios.ai/local-ai/weschera--glm52-quanttrio-int4-int8-mtp5-4x-dgx-spark.md): 32.5 tok/s (Peak; Peak, 1 request, best value reported by source, context 200k) — [source](https://github.com/Weschera/GLM-5.2-QuantTrio-4x-DGX-Spark/blob/main/README.md)
- [Qwen3.8-2.4T-A95B UD-Q1_0 with llama.cpp on 4× DGX Spark](https://www.contextstudios.ai/local-ai/weschera--qwen38-24t-a95b-ud-q1_0-mtp3-4x-dgx-spark.md): 7.76 tok/s (Peak; Peak, 1 request, best value reported by source, prompt 121 tokens) — [source](https://github.com/Weschera/Qwen3.8-2.4T-A95B-UD-Q1_0-4x-DGX-Spark/blob/main/README.md)
- [Qwen3.6-35B NVFP4 with vLLM on 1× DGX Spark](https://www.contextstudios.ai/local-ai/weschera--qwen36-35b-a3b-nvfp4-dspark8-1x-dgx-spark.md): no measured speed on record

## Related

- [Local AI on your hardware.](https://www.contextstudios.ai/local-ai.md)
