Windows prebuilt of llama.cpp combining Multi-Token Prediction (MTP) + TurboQuant KV cache compression + native sm_120 (Blackwell consumer GPU, FP4 tensor cores). For RTX 5060 Ti / 5070 / 5080 / 5090.
-
Updated
Sep 17, 2026
Windows prebuilt of llama.cpp combining Multi-Token Prediction (MTP) + TurboQuant KV cache compression + native sm_120 (Blackwell consumer GPU, FP4 tensor cores). For RTX 5060 Ti / 5070 / 5080 / 5090.
NInfer fork for 2x RTX 5060 Ti (16 GiB each): --tp 2 tensor parallelism brings the 27B package to two consumer cards. KV-cache tiers (bf16/int8/fp8/k16v8), a 253,952-token single-slot context on K16V8, MTP3 with prefix reuse that actually hits, and a /health that reports engine availability.
Run llama.cpp with Multi-Token Prediction and TurboQuant on Windows using native sm_120 Blackwell support for RTX 50-series GPUs.
To associate your repository with the rtx-5060ti topic, visit your repo's landing page and select "manage topics."