Skip to content

vulkan: let PTQ1_0 use the integer-dot mat-vec on Intel Xe2 (8.3x decode on Arc B390) - #238

Merged
bri-prism merged 5 commits into
prismfrom
perf/vulkan-intel-mmvq-ptq1_0
Sep 21, 2026
Merged

bri-prism merged 5 commits into
prismfrom
perf/vulkan-intel-mmvq-ptq1_0

Conversation

@bri-prism

Copy link
Copy Markdown
Collaborator

Stacked on #188, which adds mul_mat_vecq_ptq1_0.comp. The diff below shows #188's commits too; the net new change here is the final commit (7 lines). Happy to rebase once #188 lands.

Problem

On Intel with the proprietary Windows driver, #188's PTQ1_0 integer-dot mat-vec is never dispatched, so its decode win is invisible on that hardware.

ggml_vk_should_use_mmvq() has an INTEL_XE2 early-out that whitelists Q2_0 / Q2_K / Q3_K / Q6_K, and immediately below it:

if (device->driver_id == vk::DriverId::eIntelProprietaryWindows) {
    // Intel Windows proprietary driver MMVQ performance for !Q2/Q3/Q6 is worse than fp16,
    // see ggml-org/llama.cpp#17628 and ggml-org/llama.cpp#23056
    return false;
}

PTQ1_0 is not in the whitelist, so it returns false before reaching the per-type switch and stays on the dequant mat-vec. That comment and the two issues it links predate PTQ1_0, which did not exist when the blanket opt-out was written.

Change

Add PTQ1_0 to the INTEL_XE2 whitelist. Seven lines, five of which are the comment recording the measurements.

PQ2_0 is deliberately not added — measured below, it gains nothing from mmvq and stays on its existing dequant path.

Measurements

Intel Arc B390 (Panther Lake Xe3 iGPU, UMA, int dot: 1, matrix cores: KHR_coopmat), Intel proprietary Windows driver, Windows 11. MinGW/ucrt64 g++, Ninja, Release. Ternary-Bonsai-2-27B-PTQ1_0.gguf, -p 512 -n 128 -ngl 99 -fa 1 -r 3, default dispatch, no environment overrides:

build tg128 pp512
#188 as-is 1.55 / 1.57 154.4 / 169.3
this branch 13.02 / 12.93 156.1 / 203.8

8.3x decode. Two rounds shown; within-run spread is ±0.05 or better on the tg128 rows.

Cross-check that this is the same effect as forcing the path by hand, rather than something else: on #188 as-is, GGML_VK_FORCE_MMVQ=1 gives 13.16 — matching this branch's default-path number. The control for that is stock prism plus the same env var, which stays at 1.62, i.e. the win comes from #188's kernel and this commit only makes it reachable.

Why PQ2_0 is excluded, same machine and models:

PQ2_0 tg128
#188 default (dequant mat-vec) 11.56
#188 + GGML_VK_FORCE_MMVQ=1 11.63

Inside run-to-run spread, so there is no case for routing it through mmvq.

Correctness

test-backend-ops test -b Vulkan0, full and unfiltered, this branch vs #188:

#188 80f71b004 this branch
OK 17434 17434
not supported 4524 4524
failing cases 3 3 (same cases)
PTQ1_0 199 OK / 0 FAIL 199 OK / 0 FAIL
PQ2_0 191 OK / 0 FAIL 191 OK / 0 FAIL

Identical. The 3 failures are the pre-existing GATED_DELTA_NET raw_gates=1 bug on prism (filed as #237), unrelated to either PR.

Worth noting: because the gate blocked mmvq on this device, #188's own test run never exercised mul_mat_vecq_ptq1_0.comp here — the 199 PTQ1_0 cases went through the dequant path. With this commit those same 199 cases route through the new shader and still pass, so this is also first coverage of that kernel on Intel.

Scope

The change is inside the existing architecture == INTEL_XE2 branch, so no other vendor or Intel architecture is affected. I have only Xe3 hardware; if the opt-out is wrong for PTQ1_0 on other Intel parts too, that needs someone with those devices to say so.

alhnesn and others added 5 commits September 18, 2026 20:36
PTQ1_0 decode on Vulkan went through the generic mul_mat_vec shader, which
decodes every trit with a per-element base-3 loop; on an RX 6750 XT that ran
at ~270 GFLOPS and limited Ternary-Bonsai-2-27B to 4.7 tok/s. This adds a
dedicated q8_1 (mmvq) shader for PTQ1_0 and wires PQ2_0 into the Vulkan
backend, which had no support for it at all.

mul_mat_vecq_ptq1_0.comp:
- two lanes per 128-element block, 4 rows per workgroup, no divergence
- trits decoded with the multiply-by-3 recurrence on two bytes at a time in
  the 16-bit halves of a dword (as the CUDA vec_dot_ptq1_0_q8_1 does)
- the level word is fed to dotPacked4x8AccSat as is: activations are
  pre-masked so the residue bytes are multiplied by zero, which removes the
  per-level extraction ops
- the ternary -1 offset is folded into the q8_1 block sums

PQ2_0: block/packed16 types, dequant shader, mul_mm loader, float and q8_1
mat-vec paths (modeled on Q2_0 with the 128-element group), get_rows.

test-backend-ops: MUL_MAT / MUL_MAT_ID cases for both types at the Bonsai
shapes (k = 1024..17408), odd row counts, batched B and n = 1..8, plus perf
cases at Bonsai shapes.

Measured on RX 6750 XT (RDNA2, AMD proprietary driver, Windows 11),
Ternary-Bonsai-2-27B: PTQ1_0 decode 4.7 -> 40.8 tok/s (39 with -fa on),
PQ2_0 CPU-only -> 38.5 tok/s, prefill 200 t/s. Perplexity through the new
path 6.3916 vs 6.3906 with the original shader; greedy output identical to
the stock build.
The PTQ1_0 q8_1 mat-vec re-ran the whole trit decode for every column of B,
so batched decode (llama-server with several slots, NUM_COLS = 2..8) scaled
poorly: on an RX 6750 XT 8 parallel sequences gave 66 tok/s aggregate, barely
more than 4.

Columns are now handled in passes of CG = 3 (a new specialization constant).
Per block the lane decodes all NUM_ROWS rows level by level once and dots each
packed level with the activations of every column of the pass, so an extra
column costs three loads and three dot products per row and level. Integer
accumulators are folded into the fp32 result whenever the level moves on to a
new q8_1 sub-block scale. Passes are separate loops over the blocks, written
as guarded calls, because the driver hoists the loads of everything in one
block body (which spilled at 8 columns) and does not unroll a loop that
contains the block loop.

The single-column pipeline keeps the original loop (runtime-bounded row loop,
72 VGPRs), so single-sequence decode is unchanged.

RX 6750 XT, Ternary Bonsai 2 27B, llama-batched-bench pp256/tg128, fa on,
q8_0 KV, aggregate decode tok/s (before -> after):
  1 seq  36.6 -> 38.1     4 seq  55.2 -> 73.3
  2 seq  39.7 -> 60.8     8 seq  66.3 -> 86.4
llama-bench tg64: 40.6 (fa off) / 38.7 (fa on), same as before.
Kernel time m=17408 k=5120: n=1 45 us, n=2 57, n=4 120, n=8 215.

test-backend-ops: 271/271 (ptq1_0 + pq2_0, forced mmvq, n = 1..8 incl. row
tail), 138/138 default path; perplexity through the 8-column path 6.2321 vs
6.2294 before (same text, fp32 summation order).

Also adds perf cases with n = 2, 4, 8 for PTQ1_0 / PQ2_0 / Q4_0.
…ments

- ggml_vk_get_dequantize_mul_mat_vec_id: add PTQ1_0 and PQ2_0 to the q8_1
  whitelist. The id q8_1 pipelines were created but never selected, so
  MUL_MAT_ID always fell back to the float path.
- test-backend-ops: add n = 4 to the PTQ1_0/PQ2_0 cases (a three-column pass
  followed by the single-column tail).
- Shorten the shader comments and the coopmat2 note in vulkan-shaders-gen.cpp.

test-backend-ops MUL_MAT/MUL_MAT_ID/GET_ROWS for ptq1_0 + pq2_0: 284/284 on
the default path and 284/284 with GGML_VK_FORCE_MMVQ=1 (RX 6750 XT); the
mul_mat_vec_id_ptq1_0_q8_1_f32 pipeline is now compiled and used.
ggml_vk_should_use_mmvq() returns false for the Intel proprietary
Windows driver before reaching the per-type switch, so PTQ1_0 never
reaches mul_mat_vecq_ptq1_0.comp. Add it to the INTEL_XE2 whitelist.

Measured on Arc B390, Bonsai 2 27B, tg128: 1.56 -> 13.02 t/s.
PQ2_0 is not added: measured 11.56 -> 11.63, no gain over its
dequant mat-vec path.
@bri-prism
bri-prism merged commit 3ae4f51 into prism Sep 21, 2026
8 checks passed
@khosravipasha
khosravipasha deleted the perf/vulkan-intel-mmvq-ptq1_0 branch September 24, 2026 07:27
Preygle added a commit to Preygle/llama.cpp that referenced this pull request Sep 24, 2026
PTQ1_0 has no float mat-vec shader of its own, so it falls back to the generic
mul_mat_vec.comp, which reaches it through dequantize()/dequantize4() and calls
ptq1_0_trit() once per weight: an 8-bit load plus a loop of up to four
multiplies to skip the earlier trits in the same byte. Every packed byte is
re-read and re-decoded five times.

Since PrismML-Eng#238 an integer-dot mat-vec (mul_mat_vecq_ptq1_0.comp) covers devices
with VK_KHR_shader_integer_dot_product, and ggml_vk_should_use_mmvq() picks it
for k >= 2048 on AMD and NVIDIA, so on those GPUs this shader is not on the hot
path. It still matters for devices without integer dot support, for k < 2048,
and whenever MMVQ is declined or disabled.

mul_mat_vec_ptq1_0.comp reads each 28-byte block as seven 32-bit words and peels
the five trits off a byte with the base-3 recurrence two bytes at a time in
16-bit lanes (255*3 = 765 stays inside a lane), which is 10 multiplies per 20
weights instead of about 60. Consecutive bytes of a word map to consecutive
elements for a fixed trit, so each trit step consumes one vec4 of activations.
(trit - 1) is split into sum(trit*y) - sum(y) so the row-independent term is
gathered once per work item. A work item is half a block so rows of 5120 still
spread across the workgroup, and the loop order (word, row, column) keeps the
register footprint independent of NUM_COLS.

Measured on an RX 6700M (RDNA2, Windows, Vulkan 1.4.357) with
Ternary-Bonsai-2-27B-PTQ1_0, -ngl 99 -fa 1 -ctk q4_0 -ctv q4_0, against an
unmodified build of the parent commit with the same toolchain, runs alternated
to cancel thermal drift:

  GGML_VK_DISABLE_MMVQ=1 (this shader on the hot path)
    tg64    3.50 -> 10.44 t/s   (2.98x)
  default (MMVQ active, this shader off the hot path)
    tg128   24.6 -> 24.8 t/s    (unchanged, within noise)
    pp512   71.0 -> 70.4 t/s    (mul_mm untouched, within noise)

test-backend-ops -b Vulkan0: MUL_MAT 28/28 and MUL_MAT_ID 75/75 pass for
type_a=ptq1_0; the full suite is 17187/17190 and the 3 GATED_DELTA_NET failures
also fail on an unmodified build (36/39 both). Perplexity over a fixed
512-token sample is identical to the unmodified build: 8.6924 +/- 1.43177.
Preygle added a commit to Preygle/llama.cpp that referenced this pull request Sep 25, 2026
PTQ1_0 has no float mat-vec shader of its own, so it falls back to the generic
mul_mat_vec.comp, which reaches it through dequantize()/dequantize4() and calls
ptq1_0_trit() once per weight: an 8-bit load plus a loop of up to four
multiplies to skip the earlier trits in the same byte. Every packed byte is
re-read and re-decoded five times.

Since PrismML-Eng#238 an integer-dot mat-vec (mul_mat_vecq_ptq1_0.comp) covers devices
with VK_KHR_shader_integer_dot_product, and ggml_vk_should_use_mmvq() picks it
for k >= 2048 on AMD and NVIDIA, so on those GPUs this shader is not on the hot
path. It still matters for devices without integer dot support, for k < 2048,
and whenever MMVQ is declined or disabled.

mul_mat_vec_ptq1_0.comp reads each 28-byte block as seven 32-bit words and peels
the five trits off a byte with the base-3 recurrence two bytes at a time in
16-bit lanes (255*3 = 765 stays inside a lane), which is 10 multiplies per 20
weights instead of about 60. Consecutive bytes of a word map to consecutive
elements for a fixed trit, so each trit step consumes one vec4 of activations.
(trit - 1) is split into sum(trit*y) - sum(y) so the row-independent term is
gathered once per work item. A work item is half a block so rows of 5120 still
spread across the workgroup, and the loop order (word, row, column) keeps the
register footprint independent of NUM_COLS.

Measured on an RX 6700M (RDNA2, Windows, Vulkan 1.4.357) with
Ternary-Bonsai-2-27B-PTQ1_0, -ngl 99 -fa 1 -ctk q4_0 -ctv q4_0, against an
unmodified build of the parent commit with the same toolchain, runs alternated
to cancel thermal drift:

  GGML_VK_DISABLE_MMVQ=1 (this shader on the hot path)
    tg64    3.50 -> 10.44 t/s   (2.98x)
  default (MMVQ active, this shader off the hot path)
    tg128   24.6 -> 24.8 t/s    (unchanged, within noise)
    pp512   71.0 -> 70.4 t/s    (mul_mm untouched, within noise)

test-backend-ops -b Vulkan0: MUL_MAT 28/28 and MUL_MAT_ID 75/75 pass for
type_a=ptq1_0; the full suite is 17187/17190 and the 3 GATED_DELTA_NET failures
also fail on an unmodified build (36/39 both). Perplexity over a fixed
512-token sample is identical to the unmodified build: 8.6924 +/- 1.43177.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants