Skip to content
#

cpu-offloading

Here are 2 public repositories matching this topic...

Language: All
Filter by language

Run Qwen3.8 Flash Next on dual RTX 3090 (2x24GB) + 128GB RAM: 256K context, 75.6 tok/s at 258K input + 4K output (3 runs), 135 tok/s selected warm short-prompt decode. Pinned vLLM, CPU offload, MTP, P2P guide. Dual RTX 4090 testing notes (unvalidated).

  • Updated Sep 5, 2026
  • Python

Add this topic to your repo

To associate your repository with the cpu-offloading topic, visit your repo's landing page and select "manage topics."

Learn more