Skip to content
#

256k-context

Here is 1 public repository matching this topic...

Run Qwen3.8 Flash Next on dual RTX 3090 (2x24GB) + 128GB RAM: 256K context, 75.6 tok/s at 258K input + 4K output (3 runs), 135 tok/s selected warm short-prompt decode. Pinned vLLM, CPU offload, MTP, P2P guide. Dual RTX 4090 testing notes (unvalidated).

  • Updated Sep 5, 2026
  • Python

Add this topic to your repo

To associate your repository with the 256k-context topic, visit your repo's landing page and select "manage topics."

Learn more