Skip to content

ggml-cpu : convert each mul_mat src1 row on a single thread (faster token generation) - #249

Open
lenny76 wants to merge 1 commit into
PrismML-Eng:prismfrom
lenny76:cpu-mul-mat-src1-rowwise
Open

lenny76 wants to merge 1 commit into
PrismML-Eng:prismfrom
lenny76:cpu-mul-mat-src1-rowwise

Conversation

@lenny76

@lenny76 lenny76 commented Sep 23, 2026

Copy link
Copy Markdown

Overview

In ggml_compute_forward_mul_mat, the conversion of src1 to vec_dot_type splits every row into nth block ranges (added in ggml-org#11666 so that a single row is also converted in parallel). During token generation there is one row, so each thread writes a few blocks (with 48 threads and n = 5120: about 3 q8_0 blocks, ~100 bytes). The cache lines at the range borders are written by two cores. After the barrier, every vec_dot call that reads this row is slow, and it stays slow for the whole matmul.

This PR converts each src1 row on one thread, with the rows over (i11, i12, i13) split across threads. The repack path (ggml::cpu::repack::tensor_traits::forward_mul_mat) already does this for single rows. The converted bytes are the same, so results are bit-identical.

Additional information

Hardware: 2x Xeon Gold 6262 (Cascade Lake, 24 cores per socket, 48 cores, no SMT in the container), Linux, GCC 14, -DGGML_NATIVE=ON, CPU only. Models: Ternary-Bonsai-2-27B PTQ1_0 (vec_dot path) and PQ2_0 with --no-repack (vec_dot path, Q8_K activations). Base is prism + #245 + #246 + an AVX-512 PTQ1_0 vec_dot (separate PR). Runs alternated base/PR, two rounds.

Cost of one PTQ1_0 row dot (~1.3 KB of weights), measured with rdtsc inside mul_mat_one_chunk, 48 threads:

src1 conversion cycles per vec_dot call same call again (row already in cache)
split by blocks (master) 10500 - 12900 9500 - 11800
split by blocks aligned to 64-byte lines 4100 - 5200 3200 - 4400
one thread per row (this PR) 2600 - 3300 1900 - 2400
1 thread total 830 855

Token generation (t/s), llama-bench -p 64 -n 32, PQ2_0 via llama-completion --no-repack -n 32:

config model threads master PR
both sockets PTQ1_0 24 2.33 / 2.36 4.26 / 4.31
both sockets PTQ1_0 48 1.36 / 1.39 3.95 / 4.06
both sockets PQ2_0 no-repack 12 2.20 / 2.21 2.83 / 2.83
both sockets PQ2_0 no-repack 48 1.64 / 1.62 3.05 / 3.06
one socket (numactl -N 0 -m 0) PTQ1_0 8 2.28 / 2.27 2.73 / 2.70
one socket PTQ1_0 24 2.44 / 2.47 3.98 / 3.95
one socket PQ2_0 no-repack 8 2.88 / 2.87 3.43 / 3.42
one socket PQ2_0 no-repack 24 2.54 / 2.63 3.81 / 3.80

Prompt processing is unchanged within noise (PTQ1_0 pp64 one socket: 3.94 / 3.62 -> 3.78 / 3.81 at 8 threads, 8.97 / 8.83 -> 9.00 / 8.96 at 24; pp512 both sockets: 9.69 -> 9.83 at 24, 15.62 -> 15.90 at 48). PQ2_0 with repack is not affected by this code; its decode moves 6.63 -> 6.94 t/s (24 threads) from the bf16 matmuls in the same graph.

Correctness: greedy 48-token completions identical for PTQ1_0 and PQ2_0.

Not tested: CPUs with heterogeneous cores (the target of ggml-org#11666) and ARM. mul_mat_id has the same block split; this PR does not touch it because I have no MoE model to measure it with.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES. Claude Code (Anthropic) found the cause (per-call rdtsc timing, copy experiments), wrote the patch and ran the measurements above on my machine, under my direction.

🤖 Generated with Claude Code

The src1 conversion to vec_dot_type split every row into nth block
ranges. With one row (token generation) every thread writes a few
blocks, cache lines at the range borders are written by two cores
(often on different sockets), and all later vec_dot reads of the row
stay slow. On a 2-socket Xeon Gold 6262 with 48 threads a PTQ1_0 row
dot took ~12000 cycles instead of ~2700.

Split the conversion over flattened (i11, i12, i13) rows instead, so
each row is written by one thread. This is also what the repack path
already does. The converted data is the same.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bri-prism

Copy link
Copy Markdown
Collaborator

Tested on an Intel laptop (single-socket, AVX2-only x86).

Setup: Intel Core Ultra X7 358H (Panther Lake, 16 cores, hybrid; AVX2 + AVX-VNNI, no AVX-512), Windows 11, MSYS2 UCRT64 GCC 16.2, -DGGML_NATIVE=ON, CPU only (-ngl 0). Base = prism @ 3b19c377d; PR merged on top. llama-bench -p 64 -n 32 -r 2 -t 16, two rounds with the variant order reversed in round 2 (values are round 1 / round 2, t/s).

  • Correctness: greedy output (latest-2B-PTQ1_0, 128 tokens, temp 0) is byte-identical to base.
  • Performance: small or no effect on a single-socket client CPU, as you'd expect given the 2-socket Xeon motivation.
model base tg32 PR tg32
Bonsai 2 27B PTQ1_0 1.05 / 1.04 1.05 / 1.05
Bonsai 2 27B PQ2_0 2.75 / 2.71 2.73 / 2.78
2B PTQ1_0 8.03 / 7.08 7.41 / 7.18

Thread sweep, 2B PTQ1_0 tg32, two runs each:

threads base PR
4 8.28, 7.63 8.23, 8.02
8 9.46, 8.20 9.52, 9.42
16 6.91, 7.06 8.08, 8.06

There's a hint of a gain at 16 threads, but the main matrix doesn't back it up, so I'd call it inconclusive here. With a fast PTQ1_0 kernel, the stack (#245 + #246 + #249 + #250) vs #250 alone gives 2B tg32 19.4 / 19.3 vs 18.3 / 18.5 at t=4 and 15.4 / 15.2 vs 14.1 / 14.1 at t=8 (+5–8%). That's the combined effect of #245/#246/#249; I didn't isolate this PR.

Side note from the sweep: on this hybrid CPU, 2B decode with the fast kernel peaks at 4 threads and falls off steeply after that (18.4 → 9.7 t/s at 16).

Tested with Claude Code.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants