The newest mixture-of-experts models (GLM-5.3-Flash, MiMo-V2.6-Flash, Qwen3.8-Flash-Next, Qwen3.6) are far bigger than a gaming GPU. This fork keeps their experts in RAM and turns the free VRAM into a live cache of the experts the model is using; the GPUs compute the cached ones and the CPU the rest, at the same time. No special switches needed.
llama-server -m GLM-5.3-Flash-GSQ-RCO-3.0bit-q4kattn.gguf
llama-cli -m model.gguf -p "hello"Decode tokens/s, single stream, temperature 0, model in RAM. "Upstream" is stock llama.cpp. Machine A: 2x RTX 3090 (PCIe 4.0 x16 + chipset x4), Ryzen 7 3700X, 125 GB DDR4-3200. Machine B, a laptop: RTX 4060 8 GB, Ryzen 9 8945HS, 32 GB LPDDR5X-6400. Machine C: no GPU, Ryzen 5 3600, 64 GB DDR4-3200. GLM 3.5-bit, IQ3_S and IQ1_M are the first run of build b11707 against fresh upstream def4d406a (no discarded run before it); the other rows are hot runs of earlier builds.
Stock side = upstream llama.cpp builds def4d406a = upstream b11325, 1 Oct (the three rows above), 05af0d2b1 = b11302, 30 Sep, 4fea119 = b11041, 18 Sep, and 836d57176 = b11381, 3 Oct (docker build) for the other rows, b5cf8ce02 = b11261, 29 Sep (Windows build) on the laptop; every run is recorded with its stock build in the run log (tools/runs.db, branch runs-db). Some older stock numbers (MiMo chat, GLM 3.5-bit, IQ3_S) were measured before the run log existed and cannot be traced to a build there.
Upstream itself got faster between these builds: on Qwen3.8-Flash-Next IQ4_XS stock decode went from 27.6 t/s (upstream b11041) to 31.0 t/s (upstream b11381, median of 5), and from 25.0 to 29.3 t/s on the 12k-token prompt; GLM-5.3 3.0-bit went from 11.9 (b11302) to 12.5 t/s. Gains in this table are therefore only comparable within rows measured against the same stock build; the 12k-prompt processing speed of stock also differs a lot between those two builds (496 vs 210 t/s, not yet explained).
| model (size) | machine | test | upstream | this fork | |
|---|---|---|---|---|---|
| GLM-5.3-Flash 3.0-bit (106 GB) | A | short chat, decode | 11.9 | 22.4 | 1.9x |
| MiMo-V2.6-Flash IQ3_XXS (132 GB, bigger than RAM) | A | short chat, decode | 4.6 | 10.9 | 2.4x |
| 12k-token prompt, decode | 4.2 | 9.1 | 2.2x | ||
| 12k-token prompt, processing | 156 | 112 | 0.7x | ||
| Qwen3.8-Flash-Next UD-IQ4_XS (88 GB) * | A | short chat, decode | 27.7 | 46.5 | 1.7x |
| 12k-token prompt, decode / processing | 25.3 / 500 | 42.3 / 538 | 1.7x / 1.1x | ||
| GLM-5.3-Flash 3.5-bit (137 GB, bigger than RAM) | A | short tetris prompt, 100 tokens, decode (text-dependent) | 6.9 | 15.1 | 2.2x |
| Qwen3.8-Flash-Next GSQ IQ3_S (83 GB) | A | short tetris prompt, 100 tokens, decode | 43.9 | 57.1 | 1.3x |
| Qwen3.8-Flash-Next GSQ IQ1_M (55 GB, barely over 48 GB VRAM) | A | same | 69.1 | 67.5 (picks stock) | 1.0x |
| B | same | 11.4 | 13.5 | 1.2x | |
| B | same, prompt processing | 15.5 | 5.9 | 0.4x | |
| Qwen3.6-35B-A3B Q2_0 (11 GB, on an 8 GB GPU) | B | short tetris prompt, 100 tokens, decode | 29.2 | 59.3 | 2.0x |
| C | same, CPU only (AVX Q2_0 kernels) | 6.3 | 11.3 | 1.8x | |
| Qwen3.8-27B IQ4_NL, dense (fits VRAM) | A | same | 45.1 | 45.0 | 1.0x |
| Qwen3.8-27B IQ3_S, dense, CPU only | C | same | 1.7 | 1.6 | 1.0x |
Qwen3.8-27B Q5_K_M with MTP (--spec-type draft-mtp), fits VRAM |
A | same | 78.3 | 77.0 | 1.0x (38.6 without MTP) |
| 2.2k-token prompt, decode / processing | 31.3 / 599 | 51.1 / 792 | 1.6x / 1.3x | ||
| Models that fit in VRAM | any | anything | same | same | 1.0x (cache off) |
* from the previous release. GLM and MiMo were measured on release-candidate builds (MiMo also on b11509) before the last placement and thread commits, Qwen3.6 on the release binary. Every run behind these numbers (commit, build, machine, settings) is in
tools/bench/run-history.csv.
The defaults are chosen on your machine, not hard-coded:
- Cache or stock, measured. If the model fits in VRAM, or the cache would not pay for what you run, it is placed exactly like stock llama.cpp, so speed never drops below stock. A model more than twice your free VRAM takes the cache right away; in between, the first two runs compare both and keep the faster one.
- Self-tuning on real token times. The cache policy, the upload schedule and the decode and prompt thread counts are adjusted while you use it. A setting that does not help is dropped, a setting you fix yourself is never touched.
- It remembers. What it learned per model (hot experts, tuned settings, cache-or-stock) is kept in one file,
~/.cache/llama.cpp/moe-state.ini, so the next start, even a one-shot short prompt, begins from it. Delete the file to start over. - It tells you what it does.
llama-serverlogs, andllama-cli -lv 3prints after each reply, whether the cache is on, the hit rate, the tuned settings, threads and batch sizes.
Qwen3.8-Flash-Next MTP from PR #28243
(@danielhanchen), GLM-5.3-Flash MTP from PR
#27917 (timkhronos), both in this build. Load a model's MTP draft head
with -md:
llama-server -m Qwen3.8-Flash-Next-UD-IQ4_XS-00001-of-00003.gguf \
-md mtp-Qwen3.8-Flash-Next-shared-Q4_K_M.gguf --spec-type draft-mtp
llama-server -m GLM-5.3-Flash-GSQ-RCO-3.0bit-q4kattn.gguf \
-md GLM-5.3-Flash-MTP-Q4_K.gguf --spec-type draft-mtp --spec-draft-n-max 3- MTP heads: Qwen3.8-Flash-Next from unsloth/Qwen3.8-Flash-Next-GGUF
(
MTP/, thesharedfiles reuse the main model's embeddings); GLM-5.3-Flash from neuralll/GLM-5.3-Flash-MTP-GGUF (4.3 GiB, works with anyglm5-nextGLM-5.3-Flash GGUF). We made it: no GLM MTP GGUF existed, sotools/bench/glm_splice_mtp.pypulls just the MTP tensors out of unsloth's UD-Q4_K_XL GGUF with HTTP range requests (~4.3 GiB instead of the whole model) and writes them as a draft file; seetools/bench/to rebuild it. - GLM MTP: 89% of drafts accepted in our test, output identical to plain decoding. For a model far bigger than VRAM the draft
is not loaded unless you pass
--spec-draft-n-max.
Anything you pass is used as given and is never auto-tuned:
| you pass | effect |
|---|---|
-t N, -tb N |
fixed thread counts |
--moe-expert-cache N |
cache slots per layer; 0 turns the cache off, -1 sizes it from free VRAM |
--moe KEY=VAL,... |
cache, prefetch-slots, inserts, window, predict, train, or any tuning knob (MARGIN, GATE, WAIT, BIG, SWAP_FRAC, ...), for example --moe gate=3,margin=0 |
--load-mode pin|mmap |
pinned weights (the server default when the model fits in RAM, faster prompts) or mmap |
LLAMA_MOE_AUTO_MODE=stock|cache |
force the placement; LLAMA_MOE_STATE=0 ignores and never writes the state file |
| you pass | effect |
|---|---|
--moe cache=0 |
the whole fork off: no expert cache, nothing tuned or measured, plain stock behaviour (same as --moe-expert-cache 0) |
-at off |
only the self-tuning off (--autotune off, same as --moe autotune=0): the cache keeps working with fixed defaults, placement uses a static rule, nothing is measured or saved |
Environment forms: LLAMA_AUTOTUNE=0, LLAMA_ARG_AUTOTUNE=off, LLAMA_ARG_MOE=cache=0.
- RAM is the limit. Decode speed is bound by how fast the CPU reads the experts that are not in VRAM. More or faster RAM, more VRAM or a faster GPU link all raise it.
- Output can differ slightly from stock at temperature 0: a cached expert runs on the GPU, a missed one on the CPU.
- Prompt processing of models bigger than RAM (mmap) is 20-30% below stock: the experts stream over PCIe.
- A second GPU on a slow slot helps less; prompt processing goes to the fastest link.
The expert cache builds on @csantiago78's llama.cpp PR #27861; GLM-5.3-Flash support is upstream (#27773).
About the author of this fork: I'm actively looking for an AI engineering/research role and open to relocating out of Eastern Europe. If this work is useful to you or your team, reach out: linkedin.com/in/neuralll
LLM inference in C/C++
ggml / ops / maintainer PRs / dev stats / lib llama API / llama-server REST API
A few options to get llama.cpp installed on your machine:
- Visit https://llama.app and follow the instructions
- Run with Docker - see our Docker documentation
- Download pre-built binaries from the releases page
- Build from source by cloning this repository - check out our build guide
Once installed:
# Download and run a model directly from Hugging Face
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF
# Launch OpenAI-compatible API server
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
VLM session with llama cli
|
Built-in web UI against llama serve
|
The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on
a wide range of hardware - locally and in the cloud.
- Plain C/C++ implementation without any dependencies
- Apple silicon is a first-class citizen - optimized via ARM NEON, Accelerate and Metal frameworks
- AVX, AVX2, AVX512 and AMX support for x86 architectures
- RVV, ZVFH, ZFH, ZICBOP and ZIHINTPAUSE support for RISC-V architectures
- 1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit integer quantization for faster inference and reduced memory use
- Custom CUDA kernels for running LLMs on NVIDIA GPUs (support for AMD GPUs via HIP and Moore Threads GPUs via MUSA)
- Vulkan and SYCL backend support
- CPU+GPU hybrid inference to partially accelerate models larger than the total VRAM capacity
The llama.cpp project is build on top of the ggml library.
| Backend | Target devices |
|---|---|
| BLAS | All |
| BLIS | All |
| CANN | Ascend NPU |
| CUDA | Nvidia GPU |
| HIP | AMD GPU |
| Hexagon | Snapdragon |
| IBM zDNN | IBM Z & LinuxONE |
| MUSA | Moore Threads GPU |
| Metal | Apple Silicon |
| OpenCL | Adreno GPU |
| OpenVINO [In Progress] | Intel CPUs, GPUs, and NPUs |
| RPC | All |
| SYCL | Intel GPU |
| VirtGPU | VirtGPU APIR |
| Vulkan | GPU |
| WebGPU | All |
| ZenDNN | AMD CPU |
- How to build
- Running on Docker
- Build on Android
- Multi-GPU usage
- Performance troubleshooting
- GGML tips & tricks
- XCFramework
- Completions
- Models
- Release process
- Contributors can open PRs
- Collaborators will be invited based on contributions
- Maintainers can push to branches in the
llama.cpprepo and merge PRs into themasterbranch - Any help with managing issues, PRs and projects is very appreciated!
- Read the CONTRIBUTING.md for more information
- yhirose/cpp-httplib - Single-header HTTP server, used by
llama-server- MIT license - nothings/stb - Single-header image format decoder, used by multimodal subsystem - Public domain
- nlohmann/json - Single-header JSON library, used by various tools/examples - MIT License
- mackron/miniaudio - Single-header audio format decoder, used by multimodal subsystem - Public domain
- sheredom/subprocess.h - Single-header process launching solution for C and C++ - Public domain

