webgpu: route up to 8 columns through the mat-vec kernel - #301
Conversation
The mat-vec shader is already generic in NUM_COLS, but ggml_webgpu_mul_mat only used it for ne11 <= 4 and sent 5..8 columns to the 32x32 register tile, which pads M and runs 2-4x slower at decode-batch sizes on Apple M5 Pro (Dawn, Metal): q4_0 17408x5120 n=8 1.91 ms -> 0.99 ms q8_0 17408x5120 n=8 2.01 ms -> 1.13 ms f16 17408x5120 n=8 1.83 ms -> 0.84 ms q1_0 17408x5120 n=8 1.95 ms -> 0.87 ms n=4 unchanged; n=12/16 stay on the tile path (mat-vec regresses there). 8 matches MMVQ_MAX_BATCH_SIZE on CUDA and mul_mat_vec_max_cols on Vulkan. Also build ggml-webgpu as C++20: the Dawn webgpu_cpp.h header needs std::span and std::type_identity.
|
Added an NVIDIA H200 data point (Linux, driver 570.172.08, Dawn v20260323 prebuilt, user-space Vulkan loader). Dawn reports vendor nvidia, so this run covers the dp4a Correctness at 5 to 8 columns: us/run at 17408 x 5120, base to PR: q4_0 n=5 885 to 231, n=8 892 to 438; q8_0 n=5 843 to 244, n=8 855 to 455; f16 n=5 757 to 192, n=8 778 to 303; q1_0 n=5 699 to 221, n=8 714 to 417. n=1, 4, 12 and 16 are unchanged on both. Separately, the WebGPU tile path on this GPU is 15 to 18x slower than the CUDA backend at the same shapes (q4_0 n=12: 902 vs 56 us). Not in scope here, but worth a look before anyone relies on ggml-webgpu on NVIDIA for prefill. |
|
Intel data point, measured after the merge: Dawn runs on D3D12 here, not Vulkan (the Vulkan backend fails to load Correctness: Perf, us/run at 17408 x 5120, two interleaved rounds, before to after:
Bigger win than on Apple or the H200 because the tile path is much slower on this device: at 9 to 12 columns it is about 6.5x slower than the mat-vec kernel at n=8. A higher cap (12 or 16) is worth testing on Intel specifically; on the M5 Pro the quantized mat-vec fell off a cliff past 8, so it would need to be device-dependent. Two build issues hit along the way, both independent of this PR and applied identically to both trees:
|
What
Route matmuls with 5 to 8 columns through the WebGPU mat-vec shader instead of the 32x32 register tile, and build
ggml-webgpuas C++20.Why
ggml_webgpu_mul_matused the mat-vec kernel only forne11 <= 4. The shader is already generic inNUM_COLS, but 5 to 8 columns went to the tile path, which pads M and runs 2 to 4x slower at decode-batch sizes. Measured on an Apple M5 Pro through Dawn withtest-backend-ops perf, weight 17408 x 5120, n=8:n=4 is unchanged. Past 8 columns the mat-vec kernel regresses (register pressure on the per-column accumulators), so n > 8 stays on the tile path. The cap of 8 matches
MMVQ_MAX_BATCH_SIZEon CUDA andmul_mat_vec_max_colson Vulkan.The C++20 flag is needed because current Dawn
webgpu_cpp.husesstd::spanandstd::type_identity; without it the backend does not build against the prebuilt Dawn releases.How
One-line threshold change in
ggml_webgpu_mul_mat, plustarget_compile_features(ggml-webgpu PRIVATE cxx_std_20).Testing
test-backend-ops test -b WebGPU -o MUL_MAT -p 'n=(5|6|7|8)': 154/154 passed (f32, f16, q4_0, q4_1, q5_0, q5_1, q8_0, q1_0, iq family, mxfp4, nvfp4).mmvqmat-vec path runs. Correctness at 5 to 8 columns 154/154 on base and PR. us/run at 17408 x 5120:n=1 and n=4 are identical on both. Intel Arc (Xe3) results pending.