Repository navigation
Performance of llama.cpp with Vulkan #10879
Replies: 309 comments 477 replies
|
AMD FirePro W8100
|
|
AMD RX 470
|
|
ubuntu 24.04, vulkan and cuda installed from official APT packages.
build: 4da69d1 (4351) vs CUDA on the same build/setup
build: 4da69d1 (4351) |
|
Macbook Air M2 on Asahi Linux ggml_vulkan: Found 1 Vulkan devices:
|
|
Gentoo Linux on ROG Ally (2023) Ryzen Z1 Extreme ggml_vulkan: Found 1 Vulkan devices:
|
|
ggml_vulkan: Found 4 Vulkan devices:
|
|
build: 0d52a69 (4439) NVIDIA GeForce RTX 3090 (NVIDIA)
AMD Radeon RX 6800 XT (RADV NAVI21) (radv)
AMD Radeon (TM) Pro VII (RADV VEGA20) (radv)
Intel(R) Arc(tm) A770 Graphics (DG2) (Intel open-source Mesa driver)
|
|
@netrunnereve Some of the tg results here are a little low, I think they might be debug builds. The cmake step (at least on Linux) might require |
|
Build: 8d59d91 (4450)
Lack of proper Xe coopmat support in the ANV driver is a setback honestly.
edit: retested both with the default batch size. |
|
Here's something exotic: An AMD FirePro S10000 dual GPU from 2012 with 2x 3GB GDDR5. build: 914a82d (4452)
|
|
Latest arch with For the sake of consistency I run every bit in a script and also build every target from scratch (for some reason kill -STOP -1
timeout 240s $COMMAND
kill -CONT -1
ggml_vulkan: Found 1 Vulkan devices:
ggml_vulkan: 0 = Intel(R) Iris(R) Xe Graphics (TGL GT2) (Intel open-source Mesa driver) | uma: 1 | fp16: 1 | warp size: 32 | matrix cores: none
build: ff3fcab (4459)
This bit seems to underutilise both GPU and CPU in real conditions based on
|
|
Intel ARC A770 on Windows:
build: ba8a1f9 (4460) |
Single GPU VulkanRadeon Instinct MI25 ggml_vulkan: 0 = AMD Radeon Instinct MI25 (RADV VEGA10) (radv) | uma: 0 | fp16: 1 | warp size: 64 | matrix cores: none
build: 2739a71 (4461) Radeon PRO VII ggml_vulkan: 0 = AMD Radeon Pro VII (RADV VEGA20) (radv) | uma: 0 | fp16: 1 | warp size: 64 | matrix cores: none
build: 2739a71 (4461) Multi GPU Vulkanggml_vulkan: 0 = AMD Radeon Pro VII (RADV VEGA20) (radv) | uma: 0 | fp16: 1 | warp size: 64 | matrix cores: none
build: 2739a71 (4461) ggml_vulkan: 0 = AMD Radeon Pro VII (RADV VEGA20) (radv) | uma: 0 | fp16: 1 | warp size: 64 | matrix cores: none
build: 2739a71 (4461) Single GPU RocmDevice 0: AMD Radeon Instinct MI25, compute capability 9.0, VMM: no
build: 2739a71 (4461) Device 0: AMD Radeon Pro VII, compute capability 9.0, VMM: no
build: 2739a71 (4461) Multi GPU RocmDevice 0: AMD Radeon Pro VII, compute capability 9.0, VMM: no
build: 2739a71 (4461) Layer split
build: 2739a71 (4461) Row split
build: 2739a71 (4461) Single GPU speed is decent, but multi GPU trails Rocm by a wide margin, especially with large models due to the lack of row split. |
|
AMD Radeon RX 5700 XT on Arch using mesa-git and setting a higher GPU power limit compared to the stock card.
I also think it could be interesting adding the flash attention results to the scoreboard (even if the support for it still isn't as mature as CUDA's).
|
|
I tried but there's nothing after 1 hrs , ok, might be 40 mins... Anyway I run the llama_cli for a sample eval...
Meanwhile OpenBLAS |
|
A friend allowed me to test llama.cpp on his new laptop. It's a new Lenovo IdeaPad. Raw numbers below. Vulkan is significantly faster than CPU, especially for prompt processing. ROCm is faster for prompt processing but Vulkan is faster for token generation. This is a general trend I have observed in AMD hardware. Hardware: Software: CPU
build: cb26014 (10310) ROCm
build: cb26014 (10310) Vulkan
build: cb26014 (10310) |
|
Updated for R9700s: Minor variance between the two cards, but...significantly higher performance than last time I ran this (+10%, more or less). |
|
NVIDIA CMP 170HX (GA100, sm_80, 64 GB HBM2e) unlocked Mining card (same die gen as A100) ./bin/llama-bench -m llama-2-7b.Q4_0.gguf -ngl 99 -fa 0,1 -p 512,1024,2048,4096,8192,16384,32768 -n 128,256,512,1024,2048
build: 4c1a0af (10430) |
|
Same-machine Vulkan vs ROCm numbers on Strix Halo, for the record. Hardware: Ryzen AI MAX+ 395 / Radeon 8060S (gfx1151) / 128 GB LPDDR5X unified memory, no dedicated VRAM. Ubuntu 24.04.4, kernel 6.17. Measured UMA bandwidth ~230 GB/s. Method: llama.cpp master (2026-08), separate Vulkan (RADV) + ROCm (HIP) builds. llama-bench -p512 -n128 -b512 -t16, plus llama-server /v1/chat/completions max_tokens=200. Single request, serial, ±5%/cell. Decode throughput (tok/s):
On this UMA iGPU, Vulkan wins across the board — bare +2–25%, with MTP +15–21%. On dense 27B the bare gap is small (~3%) but MTP widens it to ~15–20%. Production pick: Vulkan + MTP. Caveat: UMA/shared-memory iGPU only — not comparable to discrete GPUs (RTX 4090 etc.). Full bilingual matrix + reproducible scripts: https://github.com/JJJJJovi/strix-halo-benchmark |
|
(deck@steamdeck llama-b10360)$ ./llama-bench -m ../../models/llama-2-7b.Q4_0.gguf -ngl 100 -fa 0,1
build: 48d22e2 (10360)
|
|
llamacpp version: b10590 llamacpp compiled with; Setting: GPU1 (Core up to 3000mz, vram up to 2650mz, power up to 350W) Setting: GPU1 + GPU0 (Core up to 3000mz, vram up to 2650mz, power up to 350W) tensor + fa 0 don't work, run with fa 1 only Very poor performance of vulkan + tensor split -> but it works and that's something. For comparison gemma 4 31B Q8: Not so bad vulkan + tensor for a larger model, it still doesn't make sense but the progress is nice looking at what was like e.g. a year ago. |
Powercolor Red Devil 7900XTXAdrenalin 26.8.1
ggml_vulkan: Found 1 Vulkan devices:
ggml_vulkan: 0 = AMD Radeon RX 7900 XTX (AMD proprietary driver) | uma: 0 | fp16: 1 | bf16: 1 | fp4: 0 | warp size: 64 | shared memory: 32768 | int dot: 1 | matrix cores: KHR_coopmat
build: 304665f (10868) Vs Nov 5th: TG Non-FA is 4.73% slower, and FA is 2.29% slower, but Non-FA PP is 5% faster and FA is 15% faster. It has been quite some time since I last benchmarked. Compared to my HIP results, Vulkan is significantly slower in PP, but still faster in TG. |
|
FYI Vulkan 1.4.359 Expands Cooperative Matrix Capabilities |
|
PS D:\works\llama-vulkan> ./llama-bench -m "D:\download\DeepSeek-R1-Distill-Llama-8B-Q4_0.gguf" -ngl 99 -fa 0,1 -p 512,1024,2048,4096,8192 -n 128,256,512,1024,2048
build: 481c65f (10903) |
|
At some point in time, Vulkan made a massive jump in prefill performance (+50%) on xe2 based iGPUs ggml_vulkan: 0 = Intel(R) Arc(TM) 140V GPU (16GB) (Intel Corporation) | uma: 1 | fp16: 1 | bf16: 0 | fp4: 0 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: KHR_coopmat
build: 4a89937 (10941) |
|
Nvidia RTX 5090 laptop (24GB, GDDR7, 256-bit) -fa 0: ggml_vulkan: 0 = NVIDIA GeForce RTX 5090 Laptop GPU (NVIDIA) | uma: 0 | fp16: 1 | bf16: 1 | fp4: 0 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: NV_coopmat2
build: 4bc272f -fa 1: ggml_vulkan: 0 = NVIDIA GeForce RTX 5090 Laptop GPU (NVIDIA) | uma: 0 | fp16: 1 | bf16: 1 | fp4: 0 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: NV_coopmat2
build: 4bc272f This laptop uses an Optimus/PRIME iGPU+dGPU combo, and requires Run at performance mode (peaked at ~160W) in Fedora workstation 44, kernel 7.2.5, Vulkan 1.4.341 Laptop: HP Omen Max 16
|
|
AMD Radeon RX 6800 (16 GB GDDR6, 256-bit), Arch Linux (kernel 7.2.5), Mesa RADV 26.2.2, Ryzen 7 5700X, Resizable BAR on. Run with ggml_vulkan: 0 = AMD Radeon RX 6800 (RADV NAVI21) (radv) | uma: 0 | fp16: dot2 | bf16: 0 | fp4: 0 | warp size: 32 | shared memory: 65536 | int dot: 1 | matrix cores: none
build: f1cee99 Compared with the current RX 6800 scoreboard entry (4b385bf), tg128 is up 6% (no FA) and 10% (FA), while pp512 is down 6–9%. Same card on the HIP backend (same commit, ROCm 7.2.4): pp512 1510 / 1739 and tg128 86.0 / 93.6 (no FA / FA). Vulkan generates faster on RDNA2, and HIP only wins on pp512 with FA on. Posted in #15021 too. Larger models on the same card (
Full write-up, including llama.cpp vs Ollama on both backends: https://claude.ai/artifact/4ALgvrvKqQnxBmECgk7bMb |
|
AMD Radeon RX 9060 XT (16 GB GDDR6, 128-bit), Arch Linux (kernel 7.2.5), Mesa RADV 26.2.2, Ryzen 5 7600X, PCIe 5.0 x16, Resizable BAR on. Run with ggml_vulkan: 0 = AMD Radeon RX 9060 XT (RADV GFX1200) (radv) | uma: 0 | fp16: dot2 | bf16: 1 | fp4: 0 | warp size: 64 | shared memory: 65536 | int dot: 1 | matrix cores: KHR_coopmat
build: f1cee99 Compared with the current RX 9060 XT scoreboard entry (ed52f36), pp512 is up 33% (no FA) and 69% (FA), and tg128 is up 1% (no FA) and 5% (FA). Same card on the HIP backend (same commit, ROCm 7.2.4): pp512 2641 / 3015 and tg128 67.3 / 70.2 (no FA / FA), so Vulkan is ahead on both. Posted in #15021 too. Larger models on the same card (
I posted RX 6800 numbers earlier in this thread with the same commit; the full RX 9060 XT vs RX 6800 comparison, including llama.cpp vs Ollama on both backends, is here: https://claude.ai/artifact/DUnSyPLJ51yYmBmou8155D |
|
A quick update on StrixHalo performance, Vulkan has improved by a lot! It is destroying ROCm now. Numbers below. Linux strixhalo 7.2.8-arch1-2 #1 SMP PREEMPT_DYNAMIC Thu, 01 Oct 2026 17:24:20 +0000 x86_64 GNU/Linux CPU
build: fb4b273 (11349) ROCM
build: fb4b273 (11349) Vulkan
build: fb4b273 (11349) |


Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
This is similar to the Apple Silicon benchmark thread, but for Vulkan! We'll be testing the Llama 2 7B model like the other thread to keep things consistent, and use Q4_0 as it's simple to compute and small enough to fit on a 4GB GPU. You can download it here.
Instructions
Either run the commands below or download one of our Vulkan releases. If you have multiple GPUs please run the test on a single GPU using
-sm none -mg YOUR_GPU_NUMBERunless the model is too big to fit in VRAM. If you use RADV please run with the environment variableRADV_PERFTEST=nogttspillas that can fix a bunch of performance issues.Share your llama-bench results along with the git hash and Vulkan info string in the comments. Feel free to try other models and compare backends, but only valid runs will be placed on the scoreboard.
If multiple entries are posted for the same setup I'll prioritize newer commits with substantial Vulkan updates, otherwise I'll pick the one with the highest overall score at my discretion. Performance may vary depending on driver, operating system, board manufacturer, etc. even if the chip is the same. For integrated graphics note that the memory speed and number of channels will greatly affect your inference speed!
Vulkan Scoreboard (Click on the headings to expand the section)
Llama 2 7B, Q4_0, no FA
Llama 2 7B, Q4_0, FA enabled
All reactions