Performance of llama.cpp on Apple Silicon M-series #4167
Replies: 97 comments 182 replies
M2 Mac Mini, 4+4 CPU, 10 GPU, 24 GB Memory (@QueryType) ✅
build: 8e672ef (1550) |
M2 Max Studio, 8+4 CPU, 38 GPU ✅
build: 8e672ef (1550) |
M2 Ultra, 16+8 CPU, 60 GPU (@crasm) ✅
build: 8e672ef (1550) |
M3 Max (MBP 16), 12+4 CPU, 40 GPU (@ymcui) ✅
build: 55978ce (1555) Short Note: mostly similar to the one reported by @slaren . But for Q4_0 |
|
How about also sharing the largest model sizes and context lengths people can run with their amount of RAM? It's important to get the amount of RAM right when buying Apple computers because you can't upgrade later. |
M2 Pro, 6+4 CPU, 16 GPU (@minosvasilias) ✅
build: e9c13ff (1560) |
|
Would love to see how M1 Max and M1 Ultra fare given their high memory bandwidth. |
M2 MAX (MBP 16) 8+4 CPU, 38 GPU, 96 GB RAM (@MrSparc) ✅
build: e9c13ff (1560) |
M1 Max (MBP 16) 8+2 CPU, 32 GPU, 64GB RAM (@CedricYauLBD) ✅
build: e9c13ff (1560) Note: M1 Max RAM Bandwidth is 400GB/s |
|
Look at what I started |
M3 Pro (MBP 14), 5+6 CPU, 14 GPU (@paramaggarwal) ✅
build: e9c13ff (1560) |
|
|
### M2 MAX (MBP 16) 38 Core 32GB ✅
build: 795cd5a (1493) |
|
I'm looking at the summary plot about "PP performance vs GPU cores" and evidence that original unquantised fp16 model always delivers more performance than quantized models. |
|
I'm holding out for a benchmark of M5 Max and assuming there is a Mac Studio launching soon. I'd like to see some numbers comparing it to a Spark and Strix Halo respectively. Anyone have $5K laying around for a new Macbook Pro ... ;) |
This comment was marked as spam.
This comment was marked as spam.
M5 Max (MBP 16"), 6+12 CPU, 40 GPU, 64 GB
build: 8e672ef (1550) |
This comment was marked as spam.
This comment was marked as spam.
Apple M-series benchmark — mac-mini-m4Hardware: Apple M4, 10 cores, 24 GB RAM, macOS 26.4 llama3.2:3bGGUF:
build: 2e97c5f (9100) qwen2.5-coder:7bGGUF:
build: 2e97c5f (9100) |
|
Apple M5 Max - 128gb - 16"
build: 52b3df0 (9754) |
|
Holy cow ... we're gonna get milked 😅
Source: Heise |
|
Instructions have been updated for the M5+ series. Let's collect the rest of the data points. |
|
M5 (base, 10-core GPU), MacBook Air 16 GB, macOS 26.6.1 ✅
build: c1d0e7a (10621) F16 does not run on this config: model weights (13.5 GiB) exceed recommendedMaxWorkingSetSize (12.7 GB), matching the blank F16 cells on other base chips. |
This comment was marked as off-topic.
This comment was marked as off-topic.
Mac mini, M5 Pro, 64 GB RAM, macOS 27.0 ✅15-core CPU (5 super + 10 performance), 16-core GPU
|
This comment was marked as off-topic.
This comment was marked as off-topic.
|
Mac Studio M5 Ultra 36/80 (256GB) ✅
|
|
Mac mini M6, 12c GPU, 32 GB RAM (BW 170GB/s), macOS 27.0.1 ✅
build: 0c1e570 |



Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Summary
LLaMA 7B
[GB/s]
Cores
[t/s]
[t/s]
[t/s]
[t/s]
[t/s]
[t/s]
plot.py
Description
This is a collection of short
llama.cppbenchmarks on various Apple Silicon hardware. It can be useful to compare the performance thatllama.cppachieves across the M-series chips and hopefully answer questions of people wondering if they should upgrade or not. Collecting info here just for Apple Silicon for simplicity. Similar collection for A-series chips is available here: #4508If you are a collaborator to the project and have an Apple Silicon device, please add your device, results and optionally username for the following command directly into this post (requires LLaMA 7B v2):
Note
These instructions have been updated as of 2026 Aug 25 to support the new M5+ chips
git checkout c1d0e7a00 cmake -B build cmake --build build --target llama-bench -j ./build/bin/llama-bench \ -hf ggml-org/Llama-2-7B-GGUF:F16 \ -hf ggml-org/Llama-2-7B-GGUF:Q8_0 \ -hf ggml-org/Llama-2-7B-GGUF:Q4_0 \ -p 512 -n 128 -ngl 99 --delay 30 2> /dev/nullPPmeans "prompt processing" (bs = 512),TGmeans "text-generation" (bs = 1),t/smeans "tokens per second"Note that in this benchmark we are evaluating the performance against the same build 8e672ef (2023 Nov 21) in order to keep all performance factors even. Since then, there have been multiple improvements resulting in better absolute performance. As an example, here is how the same test compares over time on M2 Ultra:
[GB/s]
Cores
[t/s]
[t/s]
[t/s]
[t/s]
[t/s]
[t/s]
M1 Pro, 8+2 CPU, 16 GPU (@ggerganov) ✅
build: 8e672ef (1550)
M2 Ultra, 16+8 CPU, 76 GPU (@ggerganov) ✅
build: 8e672ef (1550)
M3 Max (MBP 14), 12+4 CPU, 40 GPU (@slaren) ✅
build: d103d93 (1553)
Footnotes
https://en.wikipedia.org/wiki/Apple_M1#Variants ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8
https://en.wikipedia.org/wiki/Apple_M2#Variants ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8
https://en.wikipedia.org/wiki/Apple_M3#Variants ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8
https://en.wikipedia.org/wiki/Apple_M4#Variants ↩ ↩2 ↩3 ↩4 ↩5 ↩6
https://en.wikipedia.org/wiki/Apple_M5#Variants ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8
https://en.wikipedia.org/wiki/Apple_M6#Variants ↩ ↩2
All reactions