Conversation
Assisted-by: Devin SWE-2 max and Codex
|
Checked the "PQ2_0's float mat-vec shader itself changes on [other devices] too" part on Intel Arc B390 (Panther Lake Xe3; Vulkan, Windows 11). The B390 decodes PQ2_0 through this float mat-vec path, not MMVQ. Base Correctness: Decode,
Per kernel (
So no regression on Intel, and a small gain on the 1.7B. One review note, for non-BC-250 devices: |
|
Thanks @bri-prism for the B390 cross-check. Adding the BC-250 (gfx1013/RDNA1, RADV, no coopmat, int-dot=0) side against the same base Correctness: rebased this stack onto Scope note: on BC-250 this float mat-vec path is the only practical one (int-dot=0 makes the MMVQ/PTQ1_0 integer path fall back to generic, ~4-8 t/s vs ~29 t/s for PQ2_0), so the dedicated kernel is load-bearing here rather than net-neutral. Consistent with your B390 finding that the win is device-specific. I have decode/prefill and power numbers from a live BC-250 service, but that A/B compares an older base+stack build against a newer base+full-stack build, so I am not attributing those deltas to this PR alone. Happy to run a same-base kernel on/off comparison if useful. |
031212e to
ff6b2a0
Compare
|
Rebased onto the current BC-250 (gfx1013, RADV) on this base: |
Overview
PQ2_0 still uses the generic float mat-vec shader. On an AMD BC-250 (gfx1013/RDNA1, RADV, no cooperative matrices or accelerated integer dot), this dominates decode time. This adds a dedicated PQ2_0 shader and tunes the BC-250 path. It complements the PTQ1_0 work in #252 and extends the BC-250 tuning discussed in #229.
The shader reads each 34-byte PQ2_0 block through aligned 16-bit words, decodes each row once before looping over columns, and computes
d * (sum(q*y) - sum(y)). This algebraic reassociation changes floating-point rounding relative to the generic shader; tolerance-based operator tests pass on BC-250. The shader arithmetic matches the earlier deployed implementation used in our reference-comparison tests.BC-250 uses a 32-thread workgroup/subgroup and eight rows per workgroup (16 for five columns). For PCI
1002:13FE/ RDNA1 / no coopmat only, PQ2_0, BF16 and Q8_0 can use mat-vec through 16 columns. BF16/Q8_0 cover the supporting MTP operations; their existing shaders are reused. Other devices keep the existing launch geometry and eight-column limit, but PQ2_0's float mat-vec shader itself changes on them too.The PR excludes our packed FP16 prefill path, GDN/state-copy fusion, runtime tuning knobs, and measurement scripts.
Additional information
Measured on
prismatadfffbe41with patch commit031212ec9; since rebased onto88c4bc60b(headff6b2a06c, kernel unchanged). Both were built directly on BC-250 with GCC 16 and Vulkan enabled, Release, identical flags.Model: Ternary-Bonsai-2-27B-PQ2_0-MTP-Q8_0; MTP is not used by this llama-bench run.
-ngl 99 -fa 1 -ctk q8_0 -ctv q4_1 -b 2048 -ub 512 -t 6 -p 512 -n 128 -r 3.test-backend-ops test -b Vulkan0 -o MUL_MAT,MUL_MAT_ID -p 'type_a=(pq2_0|bf16|q8_0)'passed on both builds. Upstream: 476/476 tests passed; this PR: 514/514 tests passed. The new coverage includes columns 9 through 16, row tails, batch shapes and the 17408x5120 model shape. These are targeted operator tests. Full CI and an independent model-quality evaluation were not run on this extracted PR.Only BC-250 hardware is available for testing. The new shader also serves PQ2_0 float mat-vec on other GPUs, including the generated MUL_MAT_ID variants, so independent testing is welcome. No claim of cross-device performance parity is made.
Our earlier deployment tests compared against the previous deployed build: Korean and English/code KL top-1 agreement was 100%, with small perplexity differences. This demonstrates agreement with that reference, not an independent assessment of model quality. English long-context quality and semantic correctness of the real-request sample were not evaluated.
Requirements