Conversation
The src1 conversion to vec_dot_type split every row into nth block ranges. With one row (token generation) every thread writes a few blocks, cache lines at the range borders are written by two cores (often on different sockets), and all later vec_dot reads of the row stay slow. On a 2-socket Xeon Gold 6262 with 48 threads a PTQ1_0 row dot took ~12000 cycles instead of ~2700. Split the conversion over flattened (i11, i12, i13) rows instead, so each row is written by one thread. This is also what the repack path already does. The converted data is the same. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Tested on an Intel laptop (single-socket, AVX2-only x86). Setup: Intel Core Ultra X7 358H (Panther Lake, 16 cores, hybrid; AVX2 + AVX-VNNI, no AVX-512), Windows 11, MSYS2 UCRT64 GCC 16.2,
Thread sweep, 2B PTQ1_0 tg32, two runs each:
There's a hint of a gain at 16 threads, but the main matrix doesn't back it up, so I'd call it inconclusive here. With a fast PTQ1_0 kernel, the stack (#245 + #246 + #249 + #250) vs #250 alone gives 2B tg32 19.4 / 19.3 vs 18.3 / 18.5 at t=4 and 15.4 / 15.2 vs 14.1 / 14.1 at t=8 (+5–8%). That's the combined effect of #245/#246/#249; I didn't isolate this PR. Side note from the sweep: on this hybrid CPU, 2B decode with the fast kernel peaks at 4 threads and falls off steeply after that (18.4 → 9.7 t/s at 16). Tested with Claude Code. |
Overview
In
ggml_compute_forward_mul_mat, the conversion of src1 tovec_dot_typesplits every row intonthblock ranges (added in ggml-org#11666 so that a single row is also converted in parallel). During token generation there is one row, so each thread writes a few blocks (with 48 threads and n = 5120: about 3 q8_0 blocks, ~100 bytes). The cache lines at the range borders are written by two cores. After the barrier, everyvec_dotcall that reads this row is slow, and it stays slow for the whole matmul.This PR converts each src1 row on one thread, with the rows over (i11, i12, i13) split across threads. The repack path (
ggml::cpu::repack::tensor_traits::forward_mul_mat) already does this for single rows. The converted bytes are the same, so results are bit-identical.Additional information
Hardware: 2x Xeon Gold 6262 (Cascade Lake, 24 cores per socket, 48 cores, no SMT in the container), Linux, GCC 14,
-DGGML_NATIVE=ON, CPU only. Models:Ternary-Bonsai-2-27BPTQ1_0 (vec_dot path) and PQ2_0 with--no-repack(vec_dot path, Q8_K activations). Base isprism+ #245 + #246 + an AVX-512 PTQ1_0 vec_dot (separate PR). Runs alternated base/PR, two rounds.Cost of one PTQ1_0 row dot (~1.3 KB of weights), measured with rdtsc inside
mul_mat_one_chunk, 48 threads:Token generation (t/s),
llama-bench -p 64 -n 32, PQ2_0 viallama-completion --no-repack -n 32:numactl -N 0 -m 0)Prompt processing is unchanged within noise (PTQ1_0 pp64 one socket: 3.94 / 3.62 -> 3.78 / 3.81 at 8 threads, 8.97 / 8.83 -> 9.00 / 8.96 at 24; pp512 both sockets: 9.69 -> 9.83 at 24, 15.62 -> 15.90 at 48). PQ2_0 with repack is not affected by this code; its decode moves 6.63 -> 6.94 t/s (24 threads) from the bf16 matmuls in the same graph.
Correctness: greedy 48-token completions identical for PTQ1_0 and PQ2_0.
Not tested: CPUs with heterogeneous cores (the target of ggml-org#11666) and ARM.
mul_mat_idhas the same block split; this PR does not touch it because I have no MoE model to measure it with.Requirements
🤖 Generated with Claude Code