Conversation
PTQ1_0 decode on Vulkan went through the generic mul_mat_vec shader, which decodes every trit with a per-element base-3 loop; on an RX 6750 XT that ran at ~270 GFLOPS and limited Ternary-Bonsai-2-27B to 4.7 tok/s. This adds a dedicated q8_1 (mmvq) shader for PTQ1_0 and wires PQ2_0 into the Vulkan backend, which had no support for it at all. mul_mat_vecq_ptq1_0.comp: - two lanes per 128-element block, 4 rows per workgroup, no divergence - trits decoded with the multiply-by-3 recurrence on two bytes at a time in the 16-bit halves of a dword (as the CUDA vec_dot_ptq1_0_q8_1 does) - the level word is fed to dotPacked4x8AccSat as is: activations are pre-masked so the residue bytes are multiplied by zero, which removes the per-level extraction ops - the ternary -1 offset is folded into the q8_1 block sums PQ2_0: block/packed16 types, dequant shader, mul_mm loader, float and q8_1 mat-vec paths (modeled on Q2_0 with the 128-element group), get_rows. test-backend-ops: MUL_MAT / MUL_MAT_ID cases for both types at the Bonsai shapes (k = 1024..17408), odd row counts, batched B and n = 1..8, plus perf cases at Bonsai shapes. Measured on RX 6750 XT (RDNA2, AMD proprietary driver, Windows 11), Ternary-Bonsai-2-27B: PTQ1_0 decode 4.7 -> 40.8 tok/s (39 with -fa on), PQ2_0 CPU-only -> 38.5 tok/s, prefill 200 t/s. Perplexity through the new path 6.3916 vs 6.3906 with the original shader; greedy output identical to the stock build.
The PTQ1_0 q8_1 mat-vec re-ran the whole trit decode for every column of B, so batched decode (llama-server with several slots, NUM_COLS = 2..8) scaled poorly: on an RX 6750 XT 8 parallel sequences gave 66 tok/s aggregate, barely more than 4. Columns are now handled in passes of CG = 3 (a new specialization constant). Per block the lane decodes all NUM_ROWS rows level by level once and dots each packed level with the activations of every column of the pass, so an extra column costs three loads and three dot products per row and level. Integer accumulators are folded into the fp32 result whenever the level moves on to a new q8_1 sub-block scale. Passes are separate loops over the blocks, written as guarded calls, because the driver hoists the loads of everything in one block body (which spilled at 8 columns) and does not unroll a loop that contains the block loop. The single-column pipeline keeps the original loop (runtime-bounded row loop, 72 VGPRs), so single-sequence decode is unchanged. RX 6750 XT, Ternary Bonsai 2 27B, llama-batched-bench pp256/tg128, fa on, q8_0 KV, aggregate decode tok/s (before -> after): 1 seq 36.6 -> 38.1 4 seq 55.2 -> 73.3 2 seq 39.7 -> 60.8 8 seq 66.3 -> 86.4 llama-bench tg64: 40.6 (fa off) / 38.7 (fa on), same as before. Kernel time m=17408 k=5120: n=1 45 us, n=2 57, n=4 120, n=8 215. test-backend-ops: 271/271 (ptq1_0 + pq2_0, forced mmvq, n = 1..8 incl. row tail), 138/138 default path; perplexity through the 8-column path 6.2321 vs 6.2294 before (same text, fp32 summation order). Also adds perf cases with n = 2, 4, 8 for PTQ1_0 / PQ2_0 / Q4_0.
There was a problem hiding this comment.
🟡 Changes recommended
The new PTQ1_0 and PQ2_0 MUL_MAT_ID integer-dot pipelines are currently unreachable from the pipeline selector.
Get a fresh assessment by requesting another Copilot review.
Pull request overview
Adds optimized Vulkan PTQ1_0 integer-dot decoding and PQ2_0 backend support.
Changes:
- Adds PTQ1_0 single- and multi-column integer-dot shaders.
- Adds PQ2_0 dequantization, matmul, mat-vec, and row-access support.
- Expands correctness and performance tests.
File summaries
| File | Description |
|---|---|
| tests/test-backend-ops.cpp | Adds PTQ1_0/PQ2_0 tests and benchmarks. |
| ggml/src/ggml-vulkan/vulkan-shaders/vulkan-shaders-gen.cpp | Generates and exports new shader variants. |
| ggml/src/ggml-vulkan/vulkan-shaders/types.glsl | Defines PQ2_0 shader layouts. |
| ggml/src/ggml-vulkan/vulkan-shaders/mul_mm_funcs.glsl | Loads PQ2_0 matrix tiles. |
| ggml/src/ggml-vulkan/vulkan-shaders/mul_mat_vecq.comp | Enables PQ2_0 integer-dot dispatch. |
| ggml/src/ggml-vulkan/vulkan-shaders/mul_mat_vecq_ptq1_0.comp | Implements optimized PTQ1_0 mat-vec. |
| ggml/src/ggml-vulkan/vulkan-shaders/mul_mat_vecq_funcs.glsl | Implements PQ2_0 integer-dot helpers. |
| ggml/src/ggml-vulkan/vulkan-shaders/dequant_pq2_0.comp | Adds standalone PQ2_0 dequantization. |
| ggml/src/ggml-vulkan/vulkan-shaders/dequant_funcs.glsl | Adds shared PQ2_0 decoding helpers. |
| ggml/src/ggml-vulkan/ggml-vulkan.cpp | Registers and selects PQ2_0/PTQ1_0 pipelines. |
Review details
- Files reviewed: 10/10 changed files
- Comments generated: 4
- Review effort level: Balanced
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
|
PQ2_0 works great on Linux/RADV with an RDNA2 card. Tested at e80f16d (both commits):
|
|
Tested this on my RX 580, went from 1 token a second to roughly 6. |
…ments - ggml_vk_get_dequantize_mul_mat_vec_id: add PTQ1_0 and PQ2_0 to the q8_1 whitelist. The id q8_1 pipelines were created but never selected, so MUL_MAT_ID always fell back to the float path. - test-backend-ops: add n = 4 to the PTQ1_0/PQ2_0 cases (a three-column pass followed by the single-column tail). - Shorten the shader comments and the coopmat2 note in vulkan-shaders-gen.cpp. test-backend-ops MUL_MAT/MUL_MAT_ID/GET_ROWS for ptq1_0 + pq2_0: 284/284 on the default path and 284/284 with GGML_VK_FORCE_MMVQ=1 (RX 6750 XT); the mul_mat_vec_id_ptq1_0_q8_1_f32 pipeline is now compiled and used.
|
tested the PTQ1_0 on intel xe igpu, works great - inference 11 t/s which is more than 3x of what i get with qwen 3.8 27b q4_k_m. however, speed degrades quickly to ~6 t/s at 30K context, and prefill speed it pretty slow - starts at 80t/s empty and degrades to 30 t/s at 30K context. |
bri-prism
left a comment
There was a problem hiding this comment.
Agent review: posted by the maintainer's coding agent at their request.
No additional confirmed finding in this source pass. The current patch addresses the earlier review's missing PTQ1_0/PQ2_0 ID q8_1 whitelist entries and expands the column-count test coverage.
Vulkan compilation and numerical execution were not performed here. Retain forced-MMVQ MUL_MAT_ID coverage and all column counts 1-8, especially the three-column group plus single-column tail, as the merge gate. No independent performance claim is made by this review.
Reviewed commit: 330e828b482b208d66f83393146aaae2d68553bd.
|
Tested this PR on an AMD BC-250 (gfx1013, RADV/Mesa, Vulkan UMA ~13.6 GB; device Builds: prism base
* stock prism has no PQ2_0 Vulkan support at all (device lost on load); its Why PTQ1_0 stays slow here: the new FWIW this board also served PQ2_0 fine from llama-server at ctx 65536 with 2 Thanks for the patch — PQ2_0 on Vulkan is a big win for this board. |
|
Follow-up from the BC-250 (gfx1013, RDNA1, Two small BC-250-specific improvements, both filed as PRs stacked on this one:
Combined, the PQ2_0-MTP model now serves live on this box at ~30.2 t/s single-stream (ctx 65,536, 2 slots), zero DeviceLost/OOM. Separately, PTQ1_0 stays on the generic path on this Happy to test further builds on gfx1013. |
Data point: Intel Arc B390 (Panther Lake Xe3 iGPU), Vulkan / Windows — correctness clean, but the PTQ1_0 decode win is gated off by defaultDevice line: Intel proprietary Windows driver. This is Builds: Correctness —
|
prism 422590f5d |
#188 80f71b004 |
|
|---|---|---|
| OK | 17187 | 17434 |
| failing cases | 3 | 3 (same cases) |
| PTQ1_0 | 143 OK / 180 not-supported | 199 OK / 180 not-supported |
| PQ2_0 | 0 OK / 323 not-supported | 191 OK / 180 not-supported |
The 3 failures are GATED_DELTA_NET(... head_count=4, head_size=128, raw_gates=1), identical on stock prism, so they are pre-existing and unrelated to this PR. No new failures, and PQ2_0 goes from "no Vulkan support at all" to executing.
The finding: the new PTQ1_0 mat-vec never runs on Intel Windows by default
Bonsai 2 27B PTQ1_0, -p 512 -n 128 -ngl 99 -fa 1 -r 3, two ABAB-interleaved rounds:
| build | tg128 (r1 / r2) | pp512 (r1 / r2) |
|---|---|---|
| prism | 1.59 / 1.57 | 206.6 / 169.9 |
| #188, default | 1.55 / 1.57 | 154.4 / 169.3 |
No change. The cause is in ggml_vk_should_use_mmvq(): the INTEL_XE2 early-out whitelists only Q2_0 / Q2_K / Q3_K / Q6_K, and immediately below it
if (device->driver_id == vk::DriverId::eIntelProprietaryWindows) {
// Intel Windows proprietary driver MMVQ performance for !Q2/Q3/Q6 is worse than fp16,
return false;
}returns before the per-type switch. So on this driver PTQ1_0 (and PQ2_0) never reach mul_mat_vecq_ptq1_0.comp.
Forcing the path, same models and flags:
| build | tg128 | pp512 |
|---|---|---|
prism + GGML_VK_FORCE_MMVQ=1 |
1.62 ± 0.05 | 217.1 |
#188 + GGML_VK_FORCE_MMVQ=1 |
13.16 ± 0.01 | 215.6 |
8.1x decode, in line with the 8x in the PR description. The prism row with the same env var is the control — it stays at 1.62, so the win is this PR's kernel and not the flag.
Suggestion: let PTQ1_0 / PQ2_0 past that carve-out on Intel (e.g. add them to the INTEL_XE2 whitelist above it). The comment and the two issues it links predate both types, and on this part the blanket opt-out is costing exactly the 8x this PR is for.
PQ2_0 needs no flag
PQ2_0 goes through the dequant mat-vec this PR adds rather than mmvq, so it is unaffected by the gate: pp512 314.3, tg128 11.56 (with GGML_VK_FORCE_MMVQ=1: 294.4 / 11.63 — no meaningful difference). On stock prism this model has no Vulkan path at all.
Caveats
Single machine, and an iGPU on UMA, so "VRAM" is system DRAM at the same bandwidth. pp512 on this box carries roughly 20% round-to-round spread from page-cache state (visible in the two prism rows), so I would not read the prefill column as a regression either way; tg128 spread is under 1% and is where the signal is.
|
Followed up on the mmvq gate above with a patch rather than just a report: #238, stacked on this branch. It adds
Full Entirely your call whether to fold it into this PR or keep it separate — happy either way, and happy to rebase it if this one moves. |
|
Thanks for this work, @alhnesn. It has landed through #238, which was stacked on this branch and merged on 09-21 as |
|
Thanks @bri-prism, sorry I was away for a few days. Glad it landed through #238, and thanks for testing it on Intel and getting it working there too |
Summary
On Vulkan, PTQ1_0 decode currently goes through the generic
mul_mat_vec.comppath, which callsptq1_0_trit()per element (a base-3 loop with a branch per trit). On an RX 6750 XT that shader runs at ~270 GFLOPS and Ternary-Bonsai-2-27B decodes at 4.7 tok/s, while the same card streams Q4_0 weights at ~330 GB/s. PQ2_0 has no Vulkan support at all and falls back to the CPU.This PR adds:
mul_mat_vecq_ptq1_0.comp: a q8_1 (mmvq) shader for PTQ1_0. Two lanes share one 128-element block (lane 0: qs dwords 0..2, lane 1: dwords 3..5 + qh), 4 rows per workgroup, identical instruction stream on both lanes. Trits are decoded with the multiply-by-3 recurrence on two bytes at a time in the 16-bit halves of a dword (same idea asvec_dot_ptq1_0_q8_1in the CUDA backend). The level word is fed todotPacked4x8AccSatEXTwithout extracting the trit bits: the activation words are pre-masked so the residue bytes multiply by zero. The ternary -1 offset is folded into the q8_1 block sums.dequant_pq2_0.comp,mul_mmloader, float and q8_1 mat-vec paths (modeled on the existing Q2_0 code with the 128-element group),get_rows, and registration invulkan-shaders-gen.cpp/ggml-vulkan.cpp. Like PTQ1_0 it has no coopmat2 decoder and is skipped for cm2 generation.test-backend-opsMUL_MAT / MUL_MAT_ID cases for both types at the Bonsai shapes (k = 1024, 5120, 6144, 17408), odd row counts (m = 67 / 70), batched B, n = 1..8, and perf-mode cases at Bonsai shapes.Results (RX 6750 XT 12 GB, RDNA2, AMD proprietary driver, Windows 11, Ternary-Bonsai-2-27B)
-fa off-fa onThe PTQ1_0 kernel streams weights at ~330 GB/s (card peak 432 GB/s); the remaining per-token time is the model's other ops.
Batched decode (second commit)
The first version re-ran the trit decode for every column of B, so several sequences per step (llama-server with
-np N) scaled poorly. The second commit adds a multi-column path: columns are handled in passes of 3, and per block the lane decodes all rows once, level by level, and dots each packed level with the activations of every column of the pass. The single-column pipeline keeps the original loop unchanged.llama-batched-bench -npp 256 -ntg 128 -fa on -ctk q8_0 -ctv q8_0, aggregate decode:Kernel time for m = 17408, k = 5120: 45 µs at n = 1, 57 at n = 2, 120 at n = 4, 215 at n = 8 (before: n = 8 took about 5× n = 1). Two things worth knowing if you tune this further on other hardware: the driver hoists the loads of everything in one block body, so all passes had to be separate loops (one loop with all columns spilled at n = 8), and it does not unroll a loop that contains the block loop, which is why the passes are guarded calls. A 4-column pass was slower than 3 + 1 on this card for reasons I could not pin down (not registers).
Verification
test-backend-ops test -o MUL_MAT,MUL_MAT_ID,GET_ROWSagainst the CPU backend: all cases pass for both types, on the default paths and withGGML_VK_FORCE_MMVQ=1, with the subgroup-size workgroup and (by forcingDMMV_WG_SIZE_LARGElocally) the large-workgroup variant.types.glsl,dequant_funcs.glsl,mul_mm_funcs.glsl,mul_mat_vecq*.glsl) don't regress other formats.-ub 8): 6.3916 with the new shader vs 6.3906 with the original float path; greedy output token-identical to the stock prism-b10685 Vulkan build over 80 tokens. After the second commit the same-ub 8run (now through the multi-column path) gives 6.2321 vs 6.2294 for the first commit on a 12-chunk subset, the difference being fp32 summation order.test-backend-opscases for PTQ1_0 + PQ2_0 with forced mmvq (n = 1..8 including the row tail), 138/138 on the default path, and the pipeline statistics extension confirms no scratch use at any column count (72 VGPRs at n = 1, at most 127 at n = 8).Caveats
BLOCK_SIZEand both reduction modes, but NVIDIA / Intel / RADV runs would be very welcome.ggml_vk_should_use_mmvqis unchanged: the mmvq path is taken by the existing vendor heuristics (on AMD for k >= 2048), otherwise the float path is used as before.AI use
This kernel and the PQ2_0 port were developed using Claude Fable 5.1 in Claude Code, working interactively with the author on the target machine.