Skip to content

vulkan: add dedicated PQ2_0 mat-vec and BC-250 batch tuning - #281

Open
renovys wants to merge 1 commit into
PrismML-Eng:prismfrom
renovys:bc250-pq2-matvec
Open

renovys wants to merge 1 commit into
PrismML-Eng:prismfrom
renovys:bc250-pq2-matvec

Conversation

@renovys

@renovys renovys commented Sep 26, 2026 •

Copy link
Copy Markdown

Overview

PQ2_0 still uses the generic float mat-vec shader. On an AMD BC-250 (gfx1013/RDNA1, RADV, no cooperative matrices or accelerated integer dot), this dominates decode time. This adds a dedicated PQ2_0 shader and tunes the BC-250 path. It complements the PTQ1_0 work in #252 and extends the BC-250 tuning discussed in #229.

The shader reads each 34-byte PQ2_0 block through aligned 16-bit words, decodes each row once before looping over columns, and computes d * (sum(q*y) - sum(y)). This algebraic reassociation changes floating-point rounding relative to the generic shader; tolerance-based operator tests pass on BC-250. The shader arithmetic matches the earlier deployed implementation used in our reference-comparison tests.

BC-250 uses a 32-thread workgroup/subgroup and eight rows per workgroup (16 for five columns). For PCI 1002:13FE / RDNA1 / no coopmat only, PQ2_0, BF16 and Q8_0 can use mat-vec through 16 columns. BF16/Q8_0 cover the supporting MTP operations; their existing shaders are reused. Other devices keep the existing launch geometry and eight-column limit, but PQ2_0's float mat-vec shader itself changes on them too.

The PR excludes our packed FP16 prefill path, GDN/state-copy fusion, runtime tuning knobs, and measurement scripts.

Additional information

Measured on prism at adfffbe41 with patch commit 031212ec9; since rebased onto 88c4bc60b (head ff6b2a06c, kernel unchanged). Both were built directly on BC-250 with GCC 16 and Vulkan enabled, Release, identical flags.

llama-bench (mean of 3) upstream this PR
tg128, tokens/s 22.638 +/- 0.042 28.324 +/- 0.022 (+25.1%)
pp512, tokens/s 141.714 +/- 0.029 141.695 +/- 0.016

Model: Ternary-Bonsai-2-27B-PQ2_0-MTP-Q8_0; MTP is not used by this llama-bench run. -ngl 99 -fa 1 -ctk q8_0 -ctv q4_1 -b 2048 -ub 512 -t 6 -p 512 -n 128 -r 3.

test-backend-ops test -b Vulkan0 -o MUL_MAT,MUL_MAT_ID -p 'type_a=(pq2_0|bf16|q8_0)' passed on both builds. Upstream: 476/476 tests passed; this PR: 514/514 tests passed. The new coverage includes columns 9 through 16, row tails, batch shapes and the 17408x5120 model shape. These are targeted operator tests. Full CI and an independent model-quality evaluation were not run on this extracted PR.

Only BC-250 hardware is available for testing. The new shader also serves PQ2_0 float mat-vec on other GPUs, including the generated MUL_MAT_ID variants, so independent testing is welcome. No claim of cross-device performance parity is made.

Our earlier deployment tests compared against the previous deployed build: Korean and English/code KL top-1 agreement was 100%, with small perplexity differences. This demonstrates agreement with that reference, not an independent assessment of model quality. English long-context quality and semantic correctness of the real-request sample were not evaluated.

Requirements

  • I have read the contributing guidelines.
  • AI usage disclosure: YES. Devin SWE-2 max and Codex assisted with implementation, porting, review, benchmarking and this description; Muse max and GLM-5.3 max provided independent review. An earlier operational version also received GLM Flash assistance. The submitted patch and BC-250 test results are recorded separately from the broader deployed optimization stack.

Assisted-by: Devin SWE-2 max and Codex
@bri-prism

Copy link
Copy Markdown
Collaborator

Checked the "PQ2_0's float mat-vec shader itself changes on [other devices] too" part on Intel Arc B390 (Panther Lake Xe3; Vulkan, Windows 11). The B390 decodes PQ2_0 through this float mat-vec path, not MMVQ. Base prism @ adfffbe41 vs this PR @ 031212ec9, MinGW GCC build.

Correctness: test-backend-ops test -b Vulkan0 -o MUL_MAT -p pq2_0 gives 132/132 on base and 164/164 with the PR, including your 32 new cases.

Decode, llama-bench -ngl 99 -p 512 -n 128 -r 3, two rounds with the order flipped (t/s):

base tg128 PR tg128 base pp512 PR pp512
Ternary-Bonsai-2-27B PQ2_0 9.99 / 9.34 9.95 / 9.70 266 / 222 266 / 242
Ternary-Bonsai-1.7B PQ2_0 124.7 / 121.1 128.5 / 129.2 (~+4 %) 4173 / 4175 4222 / 3979

Per kernel (GGML_VK_PERF_LOGGER, 27B decode, MUL_MAT_VEC pq2_0, two runs): mixed on this GPU, net neutral.

  • Faster on the wide-output shapes: m=10240/12288 (k=5120), about 6–10 %.
  • Slower on the long-k shapes: m=5120 k=17408 and k=6144, about 4–8 %, and on the small m=1024 (19.3 → 20–21 µs).
  • The 248k-row output head is within noise.
  • Sum over shapes: 4755 → 4818 µs and 4601 → 4492 µs.

So no regression on Intel, and a small gain on the 1.7B.

One review note, for non-BC-250 devices: reduc_pq2 replaces the per-slot reduc for PQ2_0. On Intel (subgroups on, workgroup = subgroup size) that switches the DMMV_WG_SIZE_LARGE slot from HYBRID to SUBGROUP reduction. That's correct when the workgroup is a single subgroup, and the tests pass. It's a behaviour change beyond BC-250, though, so worth a line in the description.

@renovys

renovys commented Sep 27, 2026

Copy link
Copy Markdown
Author

Thanks @bri-prism for the B390 cross-check. Adding the BC-250 (gfx1013/RDNA1, RADV, no coopmat, int-dot=0) side against the same base prism@adfffbe41, MinGW-free native Release build.

Correctness: rebased this stack onto adfffbe41 and ran the full test-backend-ops (Vulkan0): 17673/17673 pass, no PQ2_0 regressions. The dedicated mul_mat_vec_pq2_0.comp routing and the 9..16 column-extended pipelines build and validate cleanly on the current base.

Scope note: on BC-250 this float mat-vec path is the only practical one (int-dot=0 makes the MMVQ/PTQ1_0 integer path fall back to generic, ~4-8 t/s vs ~29 t/s for PQ2_0), so the dedicated kernel is load-bearing here rather than net-neutral. Consistent with your B390 finding that the win is device-specific.

I have decode/prefill and power numbers from a live BC-250 service, but that A/B compares an older base+stack build against a newer base+full-stack build, so I am not attributing those deltas to this PR alone. Happy to run a same-base kernel on/off comparison if useful.

@renovys

renovys commented Sep 30, 2026 •

Copy link
Copy Markdown
Author

Rebased onto the current prism (88c4bc60b). The only conflict was in tests/test-backend-ops.cpp, where upstream added the PQ2_0 speculative-verify-width cases at the same spot; both blocks are kept. No kernel/shader changes versus the previous revision.

BC-250 (gfx1013, RADV) on this base: test-backend-ops 1386/1386 OK, and Korean KL/perplexity against our reference logits is identical to the previous base (same-top-p 100%). A live-server A/B against the previous base (MTP on, 2 rounds) was within noise (generation 30.62 vs 30.55 t/s, 7k prefill 162.5 vs 162.2 t/s); this build is now what our BC-250 service runs.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants