vulkan: let PTQ1_0 use the integer-dot mat-vec on Intel Xe2 (8.3x decode on Arc B390) - #238
Merged
Merged
Conversation
PTQ1_0 decode on Vulkan went through the generic mul_mat_vec shader, which decodes every trit with a per-element base-3 loop; on an RX 6750 XT that ran at ~270 GFLOPS and limited Ternary-Bonsai-2-27B to 4.7 tok/s. This adds a dedicated q8_1 (mmvq) shader for PTQ1_0 and wires PQ2_0 into the Vulkan backend, which had no support for it at all. mul_mat_vecq_ptq1_0.comp: - two lanes per 128-element block, 4 rows per workgroup, no divergence - trits decoded with the multiply-by-3 recurrence on two bytes at a time in the 16-bit halves of a dword (as the CUDA vec_dot_ptq1_0_q8_1 does) - the level word is fed to dotPacked4x8AccSat as is: activations are pre-masked so the residue bytes are multiplied by zero, which removes the per-level extraction ops - the ternary -1 offset is folded into the q8_1 block sums PQ2_0: block/packed16 types, dequant shader, mul_mm loader, float and q8_1 mat-vec paths (modeled on Q2_0 with the 128-element group), get_rows. test-backend-ops: MUL_MAT / MUL_MAT_ID cases for both types at the Bonsai shapes (k = 1024..17408), odd row counts, batched B and n = 1..8, plus perf cases at Bonsai shapes. Measured on RX 6750 XT (RDNA2, AMD proprietary driver, Windows 11), Ternary-Bonsai-2-27B: PTQ1_0 decode 4.7 -> 40.8 tok/s (39 with -fa on), PQ2_0 CPU-only -> 38.5 tok/s, prefill 200 t/s. Perplexity through the new path 6.3916 vs 6.3906 with the original shader; greedy output identical to the stock build.
The PTQ1_0 q8_1 mat-vec re-ran the whole trit decode for every column of B, so batched decode (llama-server with several slots, NUM_COLS = 2..8) scaled poorly: on an RX 6750 XT 8 parallel sequences gave 66 tok/s aggregate, barely more than 4. Columns are now handled in passes of CG = 3 (a new specialization constant). Per block the lane decodes all NUM_ROWS rows level by level once and dots each packed level with the activations of every column of the pass, so an extra column costs three loads and three dot products per row and level. Integer accumulators are folded into the fp32 result whenever the level moves on to a new q8_1 sub-block scale. Passes are separate loops over the blocks, written as guarded calls, because the driver hoists the loads of everything in one block body (which spilled at 8 columns) and does not unroll a loop that contains the block loop. The single-column pipeline keeps the original loop (runtime-bounded row loop, 72 VGPRs), so single-sequence decode is unchanged. RX 6750 XT, Ternary Bonsai 2 27B, llama-batched-bench pp256/tg128, fa on, q8_0 KV, aggregate decode tok/s (before -> after): 1 seq 36.6 -> 38.1 4 seq 55.2 -> 73.3 2 seq 39.7 -> 60.8 8 seq 66.3 -> 86.4 llama-bench tg64: 40.6 (fa off) / 38.7 (fa on), same as before. Kernel time m=17408 k=5120: n=1 45 us, n=2 57, n=4 120, n=8 215. test-backend-ops: 271/271 (ptq1_0 + pq2_0, forced mmvq, n = 1..8 incl. row tail), 138/138 default path; perplexity through the 8-column path 6.2321 vs 6.2294 before (same text, fp32 summation order). Also adds perf cases with n = 2, 4, 8 for PTQ1_0 / PQ2_0 / Q4_0.
…ments - ggml_vk_get_dequantize_mul_mat_vec_id: add PTQ1_0 and PQ2_0 to the q8_1 whitelist. The id q8_1 pipelines were created but never selected, so MUL_MAT_ID always fell back to the float path. - test-backend-ops: add n = 4 to the PTQ1_0/PQ2_0 cases (a three-column pass followed by the single-column tail). - Shorten the shader comments and the coopmat2 note in vulkan-shaders-gen.cpp. test-backend-ops MUL_MAT/MUL_MAT_ID/GET_ROWS for ptq1_0 + pq2_0: 284/284 on the default path and 284/284 with GGML_VK_FORCE_MMVQ=1 (RX 6750 XT); the mul_mat_vec_id_ptq1_0_q8_1_f32 pipeline is now compiled and used.
ggml_vk_should_use_mmvq() returns false for the Intel proprietary Windows driver before reaching the per-type switch, so PTQ1_0 never reaches mul_mat_vecq_ptq1_0.comp. Add it to the INTEL_XE2 whitelist. Measured on Arc B390, Bonsai 2 27B, tg128: 1.56 -> 13.02 t/s. PQ2_0 is not added: measured 11.56 -> 11.63, no gain over its dequant mat-vec path.
Preygle
added a commit
to Preygle/llama.cpp
that referenced
this pull request
Sep 24, 2026
PTQ1_0 has no float mat-vec shader of its own, so it falls back to the generic mul_mat_vec.comp, which reaches it through dequantize()/dequantize4() and calls ptq1_0_trit() once per weight: an 8-bit load plus a loop of up to four multiplies to skip the earlier trits in the same byte. Every packed byte is re-read and re-decoded five times. Since PrismML-Eng#238 an integer-dot mat-vec (mul_mat_vecq_ptq1_0.comp) covers devices with VK_KHR_shader_integer_dot_product, and ggml_vk_should_use_mmvq() picks it for k >= 2048 on AMD and NVIDIA, so on those GPUs this shader is not on the hot path. It still matters for devices without integer dot support, for k < 2048, and whenever MMVQ is declined or disabled. mul_mat_vec_ptq1_0.comp reads each 28-byte block as seven 32-bit words and peels the five trits off a byte with the base-3 recurrence two bytes at a time in 16-bit lanes (255*3 = 765 stays inside a lane), which is 10 multiplies per 20 weights instead of about 60. Consecutive bytes of a word map to consecutive elements for a fixed trit, so each trit step consumes one vec4 of activations. (trit - 1) is split into sum(trit*y) - sum(y) so the row-independent term is gathered once per work item. A work item is half a block so rows of 5120 still spread across the workgroup, and the loop order (word, row, column) keeps the register footprint independent of NUM_COLS. Measured on an RX 6700M (RDNA2, Windows, Vulkan 1.4.357) with Ternary-Bonsai-2-27B-PTQ1_0, -ngl 99 -fa 1 -ctk q4_0 -ctv q4_0, against an unmodified build of the parent commit with the same toolchain, runs alternated to cancel thermal drift: GGML_VK_DISABLE_MMVQ=1 (this shader on the hot path) tg64 3.50 -> 10.44 t/s (2.98x) default (MMVQ active, this shader off the hot path) tg128 24.6 -> 24.8 t/s (unchanged, within noise) pp512 71.0 -> 70.4 t/s (mul_mm untouched, within noise) test-backend-ops -b Vulkan0: MUL_MAT 28/28 and MUL_MAT_ID 75/75 pass for type_a=ptq1_0; the full suite is 17187/17190 and the 3 GATED_DELTA_NET failures also fail on an unmodified build (36/39 both). Perplexity over a fixed 512-token sample is identical to the unmodified build: 8.6924 +/- 1.43177.
This was referenced Sep 24, 2026
Preygle
added a commit
to Preygle/llama.cpp
that referenced
this pull request
Sep 25, 2026
PTQ1_0 has no float mat-vec shader of its own, so it falls back to the generic mul_mat_vec.comp, which reaches it through dequantize()/dequantize4() and calls ptq1_0_trit() once per weight: an 8-bit load plus a loop of up to four multiplies to skip the earlier trits in the same byte. Every packed byte is re-read and re-decoded five times. Since PrismML-Eng#238 an integer-dot mat-vec (mul_mat_vecq_ptq1_0.comp) covers devices with VK_KHR_shader_integer_dot_product, and ggml_vk_should_use_mmvq() picks it for k >= 2048 on AMD and NVIDIA, so on those GPUs this shader is not on the hot path. It still matters for devices without integer dot support, for k < 2048, and whenever MMVQ is declined or disabled. mul_mat_vec_ptq1_0.comp reads each 28-byte block as seven 32-bit words and peels the five trits off a byte with the base-3 recurrence two bytes at a time in 16-bit lanes (255*3 = 765 stays inside a lane), which is 10 multiplies per 20 weights instead of about 60. Consecutive bytes of a word map to consecutive elements for a fixed trit, so each trit step consumes one vec4 of activations. (trit - 1) is split into sum(trit*y) - sum(y) so the row-independent term is gathered once per work item. A work item is half a block so rows of 5120 still spread across the workgroup, and the loop order (word, row, column) keeps the register footprint independent of NUM_COLS. Measured on an RX 6700M (RDNA2, Windows, Vulkan 1.4.357) with Ternary-Bonsai-2-27B-PTQ1_0, -ngl 99 -fa 1 -ctk q4_0 -ctv q4_0, against an unmodified build of the parent commit with the same toolchain, runs alternated to cancel thermal drift: GGML_VK_DISABLE_MMVQ=1 (this shader on the hot path) tg64 3.50 -> 10.44 t/s (2.98x) default (MMVQ active, this shader off the hot path) tg128 24.6 -> 24.8 t/s (unchanged, within noise) pp512 71.0 -> 70.4 t/s (mul_mm untouched, within noise) test-backend-ops -b Vulkan0: MUL_MAT 28/28 and MUL_MAT_ID 75/75 pass for type_a=ptq1_0; the full suite is 17187/17190 and the 3 GATED_DELTA_NET failures also fail on an unmodified build (36/39 both). Perplexity over a fixed 512-token sample is identical to the unmodified build: 8.6924 +/- 1.43177.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
On Intel with the proprietary Windows driver, #188's PTQ1_0 integer-dot mat-vec is never dispatched, so its decode win is invisible on that hardware.
ggml_vk_should_use_mmvq()has anINTEL_XE2early-out that whitelistsQ2_0 / Q2_K / Q3_K / Q6_K, and immediately below it:PTQ1_0is not in the whitelist, so it returnsfalsebefore reaching the per-type switch and stays on the dequant mat-vec. That comment and the two issues it links predatePTQ1_0, which did not exist when the blanket opt-out was written.Change
Add
PTQ1_0to theINTEL_XE2whitelist. Seven lines, five of which are the comment recording the measurements.PQ2_0is deliberately not added — measured below, it gains nothing from mmvq and stays on its existing dequant path.Measurements
Intel Arc B390 (Panther Lake Xe3 iGPU, UMA,
int dot: 1,matrix cores: KHR_coopmat), Intel proprietary Windows driver, Windows 11. MinGW/ucrt64 g++, Ninja, Release.Ternary-Bonsai-2-27B-PTQ1_0.gguf,-p 512 -n 128 -ngl 99 -fa 1 -r 3, default dispatch, no environment overrides:8.3x decode. Two rounds shown; within-run spread is ±0.05 or better on the tg128 rows.
Cross-check that this is the same effect as forcing the path by hand, rather than something else: on #188 as-is,
GGML_VK_FORCE_MMVQ=1gives 13.16 — matching this branch's default-path number. The control for that is stockprismplus the same env var, which stays at 1.62, i.e. the win comes from #188's kernel and this commit only makes it reachable.Why
PQ2_0is excluded, same machine and models:GGML_VK_FORCE_MMVQ=1Inside run-to-run spread, so there is no case for routing it through mmvq.
Correctness
test-backend-ops test -b Vulkan0, full and unfiltered, this branch vs #188:80f71b004Identical. The 3 failures are the pre-existing
GATED_DELTA_NET raw_gates=1bug onprism(filed as #237), unrelated to either PR.Worth noting: because the gate blocked mmvq on this device, #188's own test run never exercised
mul_mat_vecq_ptq1_0.comphere — the 199 PTQ1_0 cases went through the dequant path. With this commit those same 199 cases route through the new shader and still pass, so this is also first coverage of that kernel on Intel.Scope
The change is inside the existing
architecture == INTEL_XE2branch, so no other vendor or Intel architecture is affected. I have only Xe3 hardware; if the opt-out is wrong forPTQ1_0on other Intel parts too, that needs someone with those devices to say so.