Skip to content

ggml-cpu : PQ2_0 AVX-512 VNNI gemm on 8-row tiles (+30-40% prompt processing) - #256

Merged
bri-prism merged 1 commit into
PrismML-Eng:prismfrom
lenny76:cpu-pq2_0-gemm-8rows
Sep 25, 2026
Merged

bri-prism merged 1 commit into
PrismML-Eng:prismfrom
lenny76:cpu-pq2_0-gemm-8rows

Conversation

@lenny76

@lenny76 lenny76 commented Sep 23, 2026

Copy link
Copy Markdown

Overview

Prompt processing with PQ2_0 on AVX-512 VNNI spends ~92% of its time in ggml_gemm_pq2_0_4x8_q8_0 (pp512 profile of Ternary-Bonsai-2-27B, FFN matmuls alone 63%). The kernel runs at about 1.4 TOPS on 24 Cascade Lake cores, around 10% of the VNNI peak. For every 4 activation rows and every 4 weight columns it expands the 2-bit weights, gathers each row out of the block_q8_0x4 interleave with a permutex2var, and recomputes dpbusd(ones, q) for the codes-1 offset, although that sum only depends on the activations.

This PR restructures the AVX-512 VNNI gemm:

  • activation rows are processed in tiles of 8 (4 for a remainder group), so each weight sub-block is expanded once per 8 rows instead of once per 4;
  • the q8_0 rows of the tile are de-interleaved once per call into a thread-local scratch, with their scales, so the inner loop does one broadcast load per row and no permutes;
  • the codes-1 correction sum_k d1_k * sum(q_k) is precomputed per row and 128-block and subtracted once per block (fnmadd with the column scale), instead of a dpbusd(ones, q) + subtract per 32-block, row and column group.

The inner loop per row and 32-block is now 2 dpbusd, 2 cvt, 2 mul, 2 fmadd. The gemv (token generation) path is unchanged.

The math is the same, but the offset is subtracted after the float scaling instead of in int32, so results are not bit-identical: the difference is summation-order noise (numbers below). A bit-exact variant (precomputed per-lane sums subtracted in int32) gave only +16% / +14%, so I did not use it.

Additional information

2x Xeon Gold 6262 (Cascade Lake), Linux, GCC 14, -DGGML_NATIVE=ON. Base prism @ 0324c66, Ternary-Bonsai-2-27B-PQ2_0.gguf.

KLD vs base (llama-perplexity -c 512 --chunks 8, 4096 tokens of Italian text): PPL ratio 1.000087 ± 0.000588, mean KLD 0.000239, max KLD 0.0026, same top token 98.97%, RMS Δp 0.33% (the same level as the summation-order noise reported for #250).

numactl --interleave=all llama-bench -ngl 0 --numa distribute -p 512 -n 32 -r 2, two rounds alternated:

threads base pp512 PR pp512 base tg32 PR tg32
24 22.93 / 23.73 32.34 / 33.48 5.47 / 5.47 5.40 / 5.49
48 36.77 / 36.90 47.03 / 47.24 5.07 / 5.09 5.30 / 5.35

With the other open CPU PRs applied (#246 #249 #250 #251 #254): pp512 26.0 -> 37.3 t/s at 24 threads and 41.5 -> 56 t/s at 48 threads.

Only the AVX-512 VNNI path changes; AVX2 and other architectures keep using their existing kernels.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES. Claude Code (Anthropic) wrote the patch and ran the measurements above on my machine, under my direction.

🤖 Generated with Claude Code

Prompt processing spent most of its time re-expanding the 2-bit
weights every 4 rows, gathering each q8_0 row out of the x4
interleave, and recomputing dpbusd(ones, q) for the codes-1 offset
per 32-block, row and column group. Rows are now processed 8 at a
time, de-interleaved once per call with their scales, and the offset
is subtracted once per 128-block from a precomputed per-row sum.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions github-actions Bot added the ggml label Sep 23, 2026
@bri-prism
bri-prism self-requested a review September 25, 2026 17:13

@bri-prism bri-prism left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tested on an AMD EPYC Genoa (AVX-512 F/BW/VL, VNNI, BF16; 8 vCPUs), CPU only, 2B PQ2_0 model. Output matches the prism base (mean KLD 0.00017, same top token 99.3%, PPL ratio 1.0002). Prompt processing pp512 goes from 111 to 143 t/s (+29%), reproduced across two alternating rounds. LGTM.

@bri-prism
bri-prism merged commit bda59ea into PrismML-Eng:prism Sep 25, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants