Run Qwen3.8 Flash Next on dual RTX 3090 (2x24GB) + 128GB RAM: 256K context, 75.6 tok/s at 258K input + 4K output (3 runs), 135 tok/s selected warm short-prompt decode. Pinned vLLM, CPU offload, MTP, P2P guide. Dual RTX 4090 testing notes (unvalidated).
moe quantization mtp multi-gpu dual-gpu rtx3090 int4 fp8 vllm local-llm llm-inference qwen speculative-decoding rtx4090 rtx-3090 cpu-offloading qwen3-8 qwen38 qwen3-8-flash-next 256k-context
-
Updated
Sep 5, 2026 - Python