Skip to content
 
 

Latest commit

 

History

12,253 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

llama.cpp fork: 1.7x to 2.4x faster decode on MoE models bigger than your VRAM

The newest mixture-of-experts models (GLM-5.3-Flash, MiMo-V2.6-Flash, Qwen3.8-Flash-Next, Qwen3.6) are far bigger than a gaming GPU. This fork keeps their experts in RAM and turns the free VRAM into a live cache of the experts the model is using; the GPUs compute the cached ones and the CPU the rest, at the same time. No special switches needed.

llama-server -m GLM-5.3-Flash-GSQ-RCO-3.0bit-q4kattn.gguf
llama-cli    -m model.gguf -p "hello"

What you get

Decode tokens/s, single stream, temperature 0, model in RAM. "Upstream" is stock llama.cpp. Machine A: 2x RTX 3090 (PCIe 4.0 x16 + chipset x4), Ryzen 7 3700X, 125 GB DDR4-3200. Machine B, a laptop: RTX 4060 8 GB, Ryzen 9 8945HS, 32 GB LPDDR5X-6400. Machine C: no GPU, Ryzen 5 3600, 64 GB DDR4-3200. GLM 3.5-bit, IQ3_S and IQ1_M are the first run of build b11707 against fresh upstream def4d406a (no discarded run before it); the other rows are hot runs of earlier builds.

Stock side = upstream llama.cpp builds def4d406a = upstream b11325, 1 Oct (the three rows above), 05af0d2b1 = b11302, 30 Sep, 4fea119 = b11041, 18 Sep, and 836d57176 = b11381, 3 Oct (docker build) for the other rows, b5cf8ce02 = b11261, 29 Sep (Windows build) on the laptop; every run is recorded with its stock build in the run log (tools/runs.db, branch runs-db). Some older stock numbers (MiMo chat, GLM 3.5-bit, IQ3_S) were measured before the run log existed and cannot be traced to a build there.

Upstream itself got faster between these builds: on Qwen3.8-Flash-Next IQ4_XS stock decode went from 27.6 t/s (upstream b11041) to 31.0 t/s (upstream b11381, median of 5), and from 25.0 to 29.3 t/s on the 12k-token prompt; GLM-5.3 3.0-bit went from 11.9 (b11302) to 12.5 t/s. Gains in this table are therefore only comparable within rows measured against the same stock build; the 12k-prompt processing speed of stock also differs a lot between those two builds (496 vs 210 t/s, not yet explained).

model (size) machine test upstream this fork
GLM-5.3-Flash 3.0-bit (106 GB) A short chat, decode 11.9 22.4 1.9x
MiMo-V2.6-Flash IQ3_XXS (132 GB, bigger than RAM) A short chat, decode 4.6 10.9 2.4x
12k-token prompt, decode 4.2 9.1 2.2x
12k-token prompt, processing 156 112 0.7x
Qwen3.8-Flash-Next UD-IQ4_XS (88 GB) * A short chat, decode 27.7 46.5 1.7x
12k-token prompt, decode / processing 25.3 / 500 42.3 / 538 1.7x / 1.1x
GLM-5.3-Flash 3.5-bit (137 GB, bigger than RAM) A short tetris prompt, 100 tokens, decode (text-dependent) 6.9 15.1 2.2x
Qwen3.8-Flash-Next GSQ IQ3_S (83 GB) A short tetris prompt, 100 tokens, decode 43.9 57.1 1.3x
Qwen3.8-Flash-Next GSQ IQ1_M (55 GB, barely over 48 GB VRAM) A same 69.1 67.5 (picks stock) 1.0x
B same 11.4 13.5 1.2x
B same, prompt processing 15.5 5.9 0.4x
Qwen3.6-35B-A3B Q2_0 (11 GB, on an 8 GB GPU) B short tetris prompt, 100 tokens, decode 29.2 59.3 2.0x
C same, CPU only (AVX Q2_0 kernels) 6.3 11.3 1.8x
Qwen3.8-27B IQ4_NL, dense (fits VRAM) A same 45.1 45.0 1.0x
Qwen3.8-27B IQ3_S, dense, CPU only C same 1.7 1.6 1.0x
Qwen3.8-27B Q5_K_M with MTP (--spec-type draft-mtp), fits VRAM A same 78.3 77.0 1.0x (38.6 without MTP)
2.2k-token prompt, decode / processing 31.3 / 599 51.1 / 792 1.6x / 1.3x
Models that fit in VRAM any anything same same 1.0x (cache off)

* from the previous release. GLM and MiMo were measured on release-candidate builds (MiMo also on b11509) before the last placement and thread commits, Qwen3.6 on the release binary. Every run behind these numbers (commit, build, machine, settings) is in tools/bench/run-history.csv.

Nothing to configure

The defaults are chosen on your machine, not hard-coded:

  • Cache or stock, measured. If the model fits in VRAM, or the cache would not pay for what you run, it is placed exactly like stock llama.cpp, so speed never drops below stock. A model more than twice your free VRAM takes the cache right away; in between, the first two runs compare both and keep the faster one.
  • Self-tuning on real token times. The cache policy, the upload schedule and the decode and prompt thread counts are adjusted while you use it. A setting that does not help is dropped, a setting you fix yourself is never touched.
  • It remembers. What it learned per model (hot experts, tuned settings, cache-or-stock) is kept in one file, ~/.cache/llama.cpp/moe-state.ini, so the next start, even a one-shot short prompt, begins from it. Delete the file to start over.
  • It tells you what it does. llama-server logs, and llama-cli -lv 3 prints after each reply, whether the cache is on, the hit rate, the tuned settings, threads and batch sizes.

MTP speculative decoding

Qwen3.8-Flash-Next MTP from PR #28243 (@danielhanchen), GLM-5.3-Flash MTP from PR #27917 (timkhronos), both in this build. Load a model's MTP draft head with -md:

llama-server -m Qwen3.8-Flash-Next-UD-IQ4_XS-00001-of-00003.gguf \
    -md mtp-Qwen3.8-Flash-Next-shared-Q4_K_M.gguf --spec-type draft-mtp
llama-server -m GLM-5.3-Flash-GSQ-RCO-3.0bit-q4kattn.gguf \
    -md GLM-5.3-Flash-MTP-Q4_K.gguf --spec-type draft-mtp --spec-draft-n-max 3
  • MTP heads: Qwen3.8-Flash-Next from unsloth/Qwen3.8-Flash-Next-GGUF (MTP/, the shared files reuse the main model's embeddings); GLM-5.3-Flash from neuralll/GLM-5.3-Flash-MTP-GGUF (4.3 GiB, works with any glm5-next GLM-5.3-Flash GGUF). We made it: no GLM MTP GGUF existed, so tools/bench/glm_splice_mtp.py pulls just the MTP tensors out of unsloth's UD-Q4_K_XL GGUF with HTTP range requests (~4.3 GiB instead of the whole model) and writes them as a draft file; see tools/bench/ to rebuild it.
  • GLM MTP: 89% of drafts accepted in our test, output identical to plain decoding. For a model far bigger than VRAM the draft is not loaded unless you pass --spec-draft-n-max.

Your settings win

Anything you pass is used as given and is never auto-tuned:

you pass effect
-t N, -tb N fixed thread counts
--moe-expert-cache N cache slots per layer; 0 turns the cache off, -1 sizes it from free VRAM
--moe KEY=VAL,... cache, prefetch-slots, inserts, window, predict, train, or any tuning knob (MARGIN, GATE, WAIT, BIG, SWAP_FRAC, ...), for example --moe gate=3,margin=0
--load-mode pin|mmap pinned weights (the server default when the model fits in RAM, faster prompts) or mmap
LLAMA_MOE_AUTO_MODE=stock|cache force the placement; LLAMA_MOE_STATE=0 ignores and never writes the state file

Turn it off

you pass effect
--moe cache=0 the whole fork off: no expert cache, nothing tuned or measured, plain stock behaviour (same as --moe-expert-cache 0)
-at off only the self-tuning off (--autotune off, same as --moe autotune=0): the cache keeps working with fixed defaults, placement uses a static rule, nothing is measured or saved

Environment forms: LLAMA_AUTOTUNE=0, LLAMA_ARG_AUTOTUNE=off, LLAMA_ARG_MOE=cache=0.

Good to know

  • RAM is the limit. Decode speed is bound by how fast the CPU reads the experts that are not in VRAM. More or faster RAM, more VRAM or a faster GPU link all raise it.
  • Output can differ slightly from stock at temperature 0: a cached expert runs on the GPU, a missed one on the CPU.
  • Prompt processing of models bigger than RAM (mmap) is 20-30% below stock: the experts stream over PCIe.
  • A second GPU on a slow slot helps less; prompt processing goes to the fastest link.

Credits and contact

The expert cache builds on @csantiago78's llama.cpp PR #27861; GLM-5.3-Flash support is upstream (#27773).

About the author of this fork: I'm actively looking for an AI engineering/research role and open to relocating out of Eastern Europe. If this work is useful to you or your team, reach out: linkedin.com/in/neuralll


llama.cpp

llama

Quick start

A few options to get llama.cpp installed on your machine:

Once installed:

# Download and run a model directly from Hugging Face
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF

# Launch OpenAI-compatible API server
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
VLM session with `llama cli` VLM session with llama cli Built-in web UI against `llama serve` running Qwen 3.6 Built-in web UI against llama serve

Description

The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on a wide range of hardware - locally and in the cloud.

  • Plain C/C++ implementation without any dependencies
  • Apple silicon is a first-class citizen - optimized via ARM NEON, Accelerate and Metal frameworks
  • AVX, AVX2, AVX512 and AMX support for x86 architectures
  • RVV, ZVFH, ZFH, ZICBOP and ZIHINTPAUSE support for RISC-V architectures
  • 1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit integer quantization for faster inference and reduced memory use
  • Custom CUDA kernels for running LLMs on NVIDIA GPUs (support for AMD GPUs via HIP and Moore Threads GPUs via MUSA)
  • Vulkan and SYCL backend support
  • CPU+GPU hybrid inference to partially accelerate models larger than the total VRAM capacity

The llama.cpp project is build on top of the ggml library.

Supported backends

Backend Target devices
BLAS All
BLIS All
CANN Ascend NPU
CUDA Nvidia GPU
HIP AMD GPU
Hexagon Snapdragon
IBM zDNN IBM Z & LinuxONE
MUSA Moore Threads GPU
Metal Apple Silicon
OpenCL Adreno GPU
OpenVINO [In Progress] Intel CPUs, GPUs, and NPUs
RPC All
SYCL Intel GPU
VirtGPU VirtGPU APIR
Vulkan GPU
WebGPU All
ZenDNN AMD CPU

Documentation

Tools

Development

Contributing

  • Contributors can open PRs
  • Collaborators will be invited based on contributions
  • Maintainers can push to branches in the llama.cpp repo and merge PRs into the master branch
  • Any help with managing issues, PRs and projects is very appreciated!
  • Read the CONTRIBUTING.md for more information

Acknowledgements

  • yhirose/cpp-httplib - Single-header HTTP server, used by llama-server - MIT license
  • nothings/stb - Single-header image format decoder, used by multimodal subsystem - Public domain
  • nlohmann/json - Single-header JSON library, used by various tools/examples - MIT License
  • mackron/miniaudio - Single-header audio format decoder, used by multimodal subsystem - Public domain
  • sheredom/subprocess.h - Single-header process launching solution for C and C++ - Public domain

About

Fork off LLama cpp that newly scales GPU 1.7x-2.4x etc also on big models not fully fitting in vram

Resources

Contributing

Security policy

Stars

10 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages