Skip to content

webgpu: route up to 8 columns through the mat-vec kernel - #301

Merged
bri-prism merged 1 commit into
prismfrom
webgpu-matvec-8cols
Oct 2, 2026
Merged

bri-prism merged 1 commit into
prismfrom
webgpu-matvec-8cols

Conversation

@bri-prism

@bri-prism bri-prism commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

What

Route matmuls with 5 to 8 columns through the WebGPU mat-vec shader instead of the 32x32 register tile, and build ggml-webgpu as C++20.

Why

ggml_webgpu_mul_mat used the mat-vec kernel only for ne11 <= 4. The shader is already generic in NUM_COLS, but 5 to 8 columns went to the tile path, which pads M and runs 2 to 4x slower at decode-batch sizes. Measured on an Apple M5 Pro through Dawn with test-backend-ops perf, weight 17408 x 5120, n=8:

type before after
q4_0 1.91 ms 0.99 ms
q8_0 2.01 ms 1.13 ms
f16 1.83 ms 0.84 ms
q1_0 1.95 ms 0.87 ms

n=4 is unchanged. Past 8 columns the mat-vec kernel regresses (register pressure on the per-column accumulators), so n > 8 stays on the tile path. The cap of 8 matches MMVQ_MAX_BATCH_SIZE on CUDA and mul_mat_vec_max_cols on Vulkan.

The C++20 flag is needed because current Dawn webgpu_cpp.h uses std::span and std::type_identity; without it the backend does not build against the prebuilt Dawn releases.

How

One-line threshold change in ggml_webgpu_mul_mat, plus target_compile_features(ggml-webgpu PRIVATE cxx_std_20).

Testing

  • test-backend-ops test -b WebGPU -o MUL_MAT -p 'n=(5|6|7|8)': 154/154 passed (f32, f16, q4_0, q4_1, q5_0, q5_1, q8_0, q1_0, iq family, mxfp4, nvfp4).
  • Perf sweep n in {4, 5, 6, 8, 12, 16} for q4_0, q8_0, f16, q1_0 on the same machine.
  • NVIDIA H200 (Linux, driver 570.172.08, Dawn v20260323 prebuilt): Dawn reports vendor nvidia, so the dp4a mmvq mat-vec path runs. Correctness at 5 to 8 columns 154/154 on base and PR. us/run at 17408 x 5120:
type n=5 n=6 n=8 n=12 (unchanged)
q4_0 885 to 231 892 to 291 892 to 438 902 / 906
q8_0 843 to 244 851 to 302 855 to 455 860 / 860
f16 757 to 192 761 to 225 778 to 303 783 / 785
q1_0 699 to 221 707 to 272 714 to 417 716 / 716

n=1 and n=4 are identical on both. Intel Arc (Xe3) results pending.

The mat-vec shader is already generic in NUM_COLS, but ggml_webgpu_mul_mat only
used it for ne11 <= 4 and sent 5..8 columns to the 32x32 register tile, which pads
M and runs 2-4x slower at decode-batch sizes on Apple M5 Pro (Dawn, Metal):

  q4_0  17408x5120  n=8   1.91 ms -> 0.99 ms
  q8_0  17408x5120  n=8   2.01 ms -> 1.13 ms
  f16   17408x5120  n=8   1.83 ms -> 0.84 ms
  q1_0  17408x5120  n=8   1.95 ms -> 0.87 ms
  n=4 unchanged; n=12/16 stay on the tile path (mat-vec regresses there).

8 matches MMVQ_MAX_BATCH_SIZE on CUDA and mul_mat_vec_max_cols on Vulkan.

Also build ggml-webgpu as C++20: the Dawn webgpu_cpp.h header needs std::span
and std::type_identity.
@bri-prism
bri-prism marked this pull request as ready for review October 2, 2026 07:36
@bri-prism
bri-prism merged commit f132654 into prism Oct 2, 2026
8 checks passed
@bri-prism

Copy link
Copy Markdown
Collaborator Author

Added an NVIDIA H200 data point (Linux, driver 570.172.08, Dawn v20260323 prebuilt, user-space Vulkan loader). Dawn reports vendor nvidia, so this run covers the dp4a mmvq mat-vec path that the Apple run could not.

Correctness at 5 to 8 columns: test-backend-ops test -b WebGPU -o MUL_MAT -p 'n=(5|6|7|8)' 154/154 on base and PR.

us/run at 17408 x 5120, base to PR: q4_0 n=5 885 to 231, n=8 892 to 438; q8_0 n=5 843 to 244, n=8 855 to 455; f16 n=5 757 to 192, n=8 778 to 303; q1_0 n=5 699 to 221, n=8 714 to 417. n=1, 4, 12 and 16 are unchanged on both.

Separately, the WebGPU tile path on this GPU is 15 to 18x slower than the CUDA backend at the same shapes (q4_0 n=12: 902 vs 56 us). Not in scope here, but worth a look before anyone relies on ggml-webgpu on NVIDIA for prefill.

@bri-prism

Copy link
Copy Markdown
Collaborator Author

Intel data point, measured after the merge: prism just before this change (f450c768f) vs just after (a14c7de99, which carries f13265492). Intel Arc B390 (Panther Lake Xe3 iGPU), Windows 11, driver 32.0.101.8724, MSVC 19.44, Ninja, Release, -DGGML_WEBGPU=ON, Dawn v20260323 prebuilt.

Dawn runs on D3D12 here, not Vulkan (the Vulkan backend fails to load vulkan-1.dll with error 87 even though it is in System32). Capabilities: vendor intel, subgroups on, dot product on, subgroup matrix off. So this run covers the vendor-gated dp4a mmvq mat-vec path on Intel, and none of the subgroup-matrix kernels.

Correctness: test-backend-ops test -b WebGPU -o MUL_MAT -p 'n=(5|6|7|8),' 162/162 before and after, 0 failures.

Perf, us/run at 17408 x 5120, two interleaved rounds, before to after:

type n=4 n=8 n=12
q4_0 505 to 498 11723 to 1540 (7.6x) 10193 to 10339
q8_0 896 to 882 11944 to 1927 (6.2x) 10450 to 10447
f16 1580 to 1570 4975 to 1600 (3.1x) 4704 to 4675
q1_0 889 to 742 (noise, one outlier round) 11898 to 1624 (7.3x) 10497 to 10376

Bigger win than on Apple or the H200 because the tile path is much slower on this device: at 9 to 12 columns it is about 6.5x slower than the mat-vec kernel at n=8. A higher cap (12 or 16) is worth testing on Intel specifically; on the M5 Pro the quantized mat-vec fell off a cliff past 8, so it would need to be device-dependent.

Two build issues hit along the way, both independent of this PR and applied identically to both trees:

  • MSVC cannot compile the generated ggml-wgsl-shaders.hpp: error C2026: string too big. embed_wgsl.py chunks raw strings at 60000 chars but MSVC's literal limit is about 16380 bytes; max_chunk_len=16000 fixes it. This blocks any MSVC build of ggml-webgpu and deserves its own small PR.
  • The Dawn package ships no DXC, so D3D12 device creation fails until dxil.dll and dxcompiler.dll from the Windows SDK sit next to the executable.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant