ggml-cpu : PQ2_0 AVX-512 VNNI gemm on 8-row tiles (+30-40% prompt processing) - #256
Merged
Merged
Conversation
Prompt processing spent most of its time re-expanding the 2-bit weights every 4 rows, gathering each q8_0 row out of the x4 interleave, and recomputing dpbusd(ones, q) for the codes-1 offset per 32-block, row and column group. Rows are now processed 8 at a time, de-interleaved once per call with their scales, and the offset is subtracted once per 128-block from a precomputed per-row sum. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
bri-prism
self-requested a review
September 25, 2026 17:13
bri-prism
reviewed
Sep 25, 2026
bri-prism
left a comment
Collaborator
There was a problem hiding this comment.
Tested on an AMD EPYC Genoa (AVX-512 F/BW/VL, VNNI, BF16; 8 vCPUs), CPU only, 2B PQ2_0 model. Output matches the prism base (mean KLD 0.00017, same top token 99.3%, PPL ratio 1.0002). Prompt processing pp512 goes from 111 to 143 t/s (+29%), reproduced across two alternating rounds. LGTM.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
Prompt processing with PQ2_0 on AVX-512 VNNI spends ~92% of its time in
ggml_gemm_pq2_0_4x8_q8_0(pp512 profile ofTernary-Bonsai-2-27B, FFN matmuls alone 63%). The kernel runs at about 1.4 TOPS on 24 Cascade Lake cores, around 10% of the VNNI peak. For every 4 activation rows and every 4 weight columns it expands the 2-bit weights, gathers each row out of theblock_q8_0x4interleave with apermutex2var, and recomputesdpbusd(ones, q)for the codes-1 offset, although that sum only depends on the activations.This PR restructures the AVX-512 VNNI gemm:
sum_k d1_k * sum(q_k)is precomputed per row and 128-block and subtracted once per block (fnmaddwith the column scale), instead of adpbusd(ones, q)+ subtract per 32-block, row and column group.The inner loop per row and 32-block is now 2
dpbusd, 2cvt, 2mul, 2fmadd. The gemv (token generation) path is unchanged.The math is the same, but the offset is subtracted after the float scaling instead of in int32, so results are not bit-identical: the difference is summation-order noise (numbers below). A bit-exact variant (precomputed per-lane sums subtracted in int32) gave only +16% / +14%, so I did not use it.
Additional information
2x Xeon Gold 6262 (Cascade Lake), Linux, GCC 14,
-DGGML_NATIVE=ON. Baseprism@0324c66,Ternary-Bonsai-2-27B-PQ2_0.gguf.KLD vs base (
llama-perplexity -c 512 --chunks 8, 4096 tokens of Italian text): PPL ratio 1.000087 ± 0.000588, mean KLD 0.000239, max KLD 0.0026, same top token 98.97%, RMS Δp 0.33% (the same level as the summation-order noise reported for #250).numactl --interleave=all llama-bench -ngl 0 --numa distribute -p 512 -n 32 -r 2, two rounds alternated:With the other open CPU PRs applied (#246 #249 #250 #251 #254): pp512 26.0 -> 37.3 t/s at 24 threads and 41.5 -> 56 t/s at 48 threads.
Only the AVX-512 VNNI path changes; AVX2 and other architectures keep using their existing kernels.
Requirements
🤖 Generated with Claude Code