---
type: "ProfilePage"
title: "Weschera"
description: "24 Rezepte von Weschera für lokale Modelle, mit Tempo, Quelle und Gewichten."
resource: "https://www.contextstudios.ai/de/lokale-ki/kreatoren/weschera"
language: "de"
generated:
  by: "process:contextstudios-md/1"
  at: "2026-10-02T21:53:20.136Z"
status: "stable"
---

# Weschera


- [Qwen3.8-27B mit MLX auf 1× Mac](https://www.contextstudios.ai/de/lokale-ki/weschera--qwen38-27b-omlx-ane-mtp3-mac-studio-m4-max.md): 53,3 tok/s (Alltag; Alltag, 1 Anfrage, Prosa-Prompt, mit MTP (nativ) (laut Rezept)) — [source](https://github.com/Weschera/Qwen3.8-27B-oMLX-MTP-Mac/blob/main/README.md)
- [Qwen3.8-27B NVFP4 mit vLLM auf 1× DGX Spark (Weschera)](https://www.contextstudios.ai/de/lokale-ki/weschera--qwen38-27b-nvfp4-unsloth-mtp3-1x-dgx-spark.md): 45 tok/s (Alltag; Alltag, 1 Anfrage, Prosa-Prompt, mit MTP) — [source](https://github.com/Weschera/Qwen3.8-27B-DGX-Spark-Quant-Ladder/blob/main/README.md)
- [Qwen3.8-Flash-Next NVFP4 mit vLLM auf 2× DGX Spark (Weschera)](https://www.contextstudios.ai/de/lokale-ki/weschera--qwen38-flash-next-nvfp4-nvidia-mtp3-2x-dgx-spark.md): 33 tok/s (Alltag; Alltag, 1 Anfrage, Prosa-Prompt, mit MTP) — [source](https://github.com/Weschera/Qwen3.8-Flash-Next-NVFP4-2x-DGX-Spark/blob/main/README.md)
- [DeepSeek-V4-Flash 284B IQ2_XXS mit DwarfStar auf 1× Mac (Vision)](https://www.contextstudios.ai/de/lokale-ki/weschera--dsv4-flash-vision-exp-iq2xxs-dspark-ds4-mac-studio-m4-max.md): 28,5 tok/s (Alltag; Alltag, 1 Anfrage, Prosa-Prompt, Kontext 32k, mit DSpark (laut Rezept)) — [source](https://github.com/Weschera/DeepSeek-V4-Flash-Vision-Exp-ds4-Mac-Studio/blob/main/README.md)
- [Ling-3.0-flash-VL FP8 mit SGLang auf 2× DGX Spark](https://www.contextstudios.ai/de/lokale-ki/weschera--ling30-flash-vl-fp8-sglang-tp2-2x-dgx-spark.md): 26 tok/s (Alltag; Alltag, 1 Anfrage, Prosa-Prompt, ohne Spekulation) — [source](https://github.com/Weschera/Ling-3.0-flash-VL-2x-DGX-Sparks/blob/main/README.md)
- [Qwen3.8-Flash-Next 125B UD-Q4_K_XL mit llama.cpp auf 1× DGX Spark](https://www.contextstudios.ai/de/lokale-ki/weschera--qwen38-flash-next-ud-q4kxl-mtp4-1x-dgx-spark.md): 24,2 tok/s (Alltag; Alltag, 1 Anfrage, Prosa-Prompt, mit Draft) — [source](https://github.com/Weschera/Qwen3.8-Flash-Next-1x-DGX-Spark/blob/main/README.md)
- [GLM-5.3-Flash NVFP4 mit vLLM auf 2× DGX Spark (128k Kontext, mit MTP)](https://www.contextstudios.ai/de/lokale-ki/weschera--glm53-flash-nvfp4-vcruz-mtp4-2x-dgx-spark.md): 23 tok/s (Alltag; Alltag, 1 Anfrage, Prosa-Prompt, mit MTP4) — [source](https://github.com/Weschera/glm53-flash-2spark-tp2/blob/main/README.md)
- [Spark-X2.5-4B BF16 mit SGLang auf 1× DGX Spark](https://www.contextstudios.ai/de/lokale-ki/weschera--spark-x25-4b-bf16-sglang-1x-dgx-spark.md): 21,8 tok/s (Alltag; Alltag, 1 Anfrage, Prosa-Prompt, ohne Spekulation) — [source](https://github.com/Weschera/spark-x25-4b-dgx-recipe/blob/main/README.md)
- [GLM-5.3-Flash NVFP4 mit vLLM auf 2× DGX Spark (128k Kontext, mit DFlash2)](https://www.contextstudios.ai/de/lokale-ki/weschera--glm53-flash-nvfp4-nvidia-dflash2-k7-2x-dgx-spark.md): 20,5 tok/s (Alltag; Alltag, 1 Anfrage, Prosa-Prompt, ohne Spekulation) — [source](https://github.com/Weschera/GLM-5.3-Flash-NVIDIA-NVFP4-2x-DGX-Spark/blob/main/README.md)
- [Laguna-S-2.1 118B NVFP4 mit vLLM auf 1× DGX Spark (128k Kontext)](https://www.contextstudios.ai/de/lokale-ki/weschera--laguna-s-21-nvfp4-dflash15-1x-dgx-spark.md): 19,3 tok/s (Alltag; Alltag, 1 Anfrage, Prosa-Prompt, Kontext 1k, ohne Spekulation) — [source](https://github.com/Weschera/Laguna-S-2.1-NVFP4-1x-DGX-Spark/blob/main/README.md)
- [Qwen3.8-27B NVFP4 mit Atlas auf 1× DGX Spark](https://www.contextstudios.ai/de/lokale-ki/weschera--qwen38-27b-nvfp4-atlas-mtp3-1x-dgx-spark.md): 19 tok/s (Alltag; Alltag, 1 Anfrage, Prosa-Prompt, mit MTP (laut Rezept)) — [source](https://github.com/Weschera/Atlas-DGX-Spark-Quickstart/blob/main/README.md)
- [GLM-5.3-Flash 320B UD-Q2_K_XL mit llama.cpp auf 1× DGX Spark](https://www.contextstudios.ai/de/lokale-ki/weschera--glm53-flash-ud-q2kxl-mtp-1x-dgx-spark.md): 17,9 tok/s (Alltag; Alltag, 1 Anfrage, Prosa-Prompt, Kontext 32k, ohne Spekulation) — [source](https://github.com/Weschera/glm53-flash-one-spark/blob/main/README.md)
- [Bonsai-2-27B Q8_0 mit llama.cpp auf 1× Mac](https://www.contextstudios.ai/de/lokale-ki/weschera--bonsai-2-27b-pq2_0-prism-llamacpp-mac-studio-m4-max.md): 10,6 tok/s (Alltag; Alltag, 1 Anfrage, Prosa-Prompt, ohne Spekulation) — [source](https://github.com/Weschera/Bonsai-2-27B-Mac/blob/main/README.md)
- [DeepSeek-V4-Flash FP8 mit vLLM auf 2× DGX Spark (1024k Kontext)](https://www.contextstudios.ai/de/lokale-ki/weschera--dsv4-flash-0731-dspark-k7-2x-dgx-spark.md): 83,8 tok/s (Gemischt; Gemischt, 1 Anfrage, gemischtes Prompt-Set, mit Draft) — [source](https://github.com/Weschera/DeepSeek-V4-Flash-0731-DSpark-2x-DGX-Spark/blob/main/README.md)
- [Qwen3.8-Flash-Next NVFP4 mit SGLang auf 2× DGX Spark](https://www.contextstudios.ai/de/lokale-ki/weschera--qwen38-flash-next-nvfp4-radixark-nextn-2x-dgx-spark.md): 41,7 tok/s (Gemischt; Gemischt, 1 Anfrage, gemischtes Prompt-Set, mit NEXTN (MTP) (laut Rezept)) — [source](https://github.com/Weschera/qwen38-flashnext-dgx-spark/blob/main/README.md)
- [Qwen3.8-Flash-Next 80B UD-IQ4_XS mit llama.cpp auf 1× DGX Spark](https://www.contextstudios.ai/de/lokale-ki/weschera--qwen38-flash-next-ud-iq4xs-262k-1x-dgx-spark.md): 26,9 tok/s (Gemischt; Gemischt, 1 Anfrage, gemischtes Prompt-Set, ohne Spekulation) — [source](https://github.com/Weschera/qwen38-flashnext-single-spark/blob/main/README.md)
- [Qwen3.8-27B NVFP4 mit SGLang auf 1× DGX Spark (Weschera)](https://www.contextstudios.ai/de/lokale-ki/weschera--qwen38-27b-nvfp4-dflash2-1x-dgx-spark.md): 25,2 tok/s (Gemischt; Gemischt, 1 Anfrage, gemischtes Prompt-Set, mit DFlash2 (laut Rezept)) — [source](https://github.com/Weschera/Qwen3.8-27B-NVFP4-DFlash2-DGX-Spark/blob/main/README.md)
- [GLM-5.3-Flash 320B UD-IQ3_XXS mit llama.cpp auf 1× DGX Spark](https://www.contextstudios.ai/de/lokale-ki/weschera--glm53-flash-ud-iq3xxs-mtp2-1x-dgx-spark.md): 20,8 tok/s (Gemischt; Gemischt, 1 Anfrage, gemischtes Prompt-Set, mit MTP) — [source](https://github.com/Weschera/GLM-5.3-Flash-Unsloth-1x-DGX-Spark/blob/main/README.md)
- [Qwen3.6-35B NVFP4 mit SGLang auf 1× DGX Spark](https://www.contextstudios.ai/de/lokale-ki/weschera--qwen36-35b-a3b-nvfp4-nextn-sglang-1x-dgx-spark.md): 92,2 tok/s (Spitze; Spitze, 1 Anfrage, Bestwert der Quelle, Kontext 64k) — [source](https://github.com/Weschera/qwen-sglang-dgx-spark/blob/main/README.md)
- [Qwen3.8-Flash-Next 125B mit MLX auf 1× Mac](https://www.contextstudios.ai/de/lokale-ki/weschera--qwen38-flash-next-omlx-mtp6-mac-studio-m4-max.md): 83,1 tok/s (Spitze; Spitze, 1 Anfrage, Code-Prompt) — [source](https://github.com/Weschera/qwen38-flash-next-omlx-mac/blob/main/README.md)
- [DeepSeek-V4-Flash NVFP4 mit vLLM auf 2× DGX Spark](https://www.contextstudios.ai/de/lokale-ki/weschera--dsv4-flash-dspark-1m-nvfp4kv-2x-dgx-spark.md): 65,4 tok/s (Spitze; Spitze, 1 Anfrage, Bestwert der Quelle) — [source](https://github.com/Weschera/DeepSeek-v4-Flash-DSpark-1M-NVFP4-KV-2x-DGX-Spark/blob/main/README.md)
- [GLM-5.2 753B INT4/INT8 mit vLLM auf 4× DGX Spark](https://www.contextstudios.ai/de/lokale-ki/weschera--glm52-quanttrio-int4-int8-mtp5-4x-dgx-spark.md): 32,5 tok/s (Spitze; Spitze, 1 Anfrage, Bestwert der Quelle, Kontext 200k) — [source](https://github.com/Weschera/GLM-5.2-QuantTrio-4x-DGX-Spark/blob/main/README.md)
- [Qwen3.8-2.4T-A95B UD-Q1_0 mit llama.cpp auf 4× DGX Spark](https://www.contextstudios.ai/de/lokale-ki/weschera--qwen38-24t-a95b-ud-q1_0-mtp3-4x-dgx-spark.md): 7,76 tok/s (Spitze; Spitze, 1 Anfrage, Bestwert der Quelle, Prompt 121 Token) — [source](https://github.com/Weschera/Qwen3.8-2.4T-A95B-UD-Q1_0-4x-DGX-Spark/blob/main/README.md)
- [Qwen3.6-35B NVFP4 mit vLLM auf 1× DGX Spark](https://www.contextstudios.ai/de/lokale-ki/weschera--qwen36-35b-a3b-nvfp4-dspark8-1x-dgx-spark.md): kein Tempowert belegt

## Related

- [Lokale KI auf Ihrer Hardware.](https://www.contextstudios.ai/de/lokale-ki.md)
