Skip to content

CUDA: Optimize accumulation in mmq for NVFP4 type - #29857

Merged
ORippler merged 3 commits into
ggml-org:masterfrom
kmorennv:optimize_mmq_nvfp4
Oct 5, 2026
Merged

ORippler merged 3 commits into
ggml-org:masterfrom
kmorennv:optimize_mmq_nvfp4

Conversation

@kmorennv

@kmorennv kmorennv commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

Overview

Simplified internal computation by accumulating results directly into the final output buffer instead of using a temporary intermediate step.

Additional information

The example to explain perf-gain use is Qwen3.6-35B-A3B-NVFP4, comparing to master 8d81559fa against this branch:

Settings: batch and microbatch 4,096, Flash Attention enabled, full GPU offload, RTX PRO 6000 Blackwell, CTK 13.4 Win64. Throughput used seven repetitions; nsys used three measured prompts after warmup.

Measurement Master Branch Difference
NVFP4 MMQ time per 8,192-token prompt, nsys 153.01 ms 142.69 ms 6.74% less time
MMQ throughput for the same work 1.00x 1.072x 7.23% higher
Full prompt processing time 564.30 ms 549.72 ms 14.58 ms saved
Full prompt throughput 14,521 tokens/s 14,907 tokens/s 2.65% higher

The kernel improvement is visible across matched launch shapes. Both builds execute 860 NVFP4 MMQ launches per prompt. For example, one matching kernel/grid falls from 808.39 us to 751.55 us, a 7.03% reduction.

The smaller overall throughput gain also makes sense: NVFP4 MMQ accounts for approximately 31% of summed GPU kernel time in master. Reducing that component by 6.74% saves roughly 2.1% of total kernel time if everything else stays constant, consistent with the observed 2.65% prompt-throughput gain.

Nsight Compute confirms the effect across 12 matched NVFP4 launches from this model:

Counter Master Branch
Executed instructions, summed 995.09 million 769.44 million
Registers per thread 255 232-236
Executed register-spill instructions, summed 17.76 million 0
Achieved occupancy, mean 16.46% 16.45%

Across the other completed nsys captures, at the same batch size:

Model Master NVFP4 MMQ Branch NVFP4 MMQ Time reduction
Qwen3.6-27B NVFP4 453.48 ms 418.46 ms 7.72%
Qwen3.6-27B NVFP4_fp8_q8 373.91 ms 352.35 ms 5.77%
Gemma-4-26B-A4B 97.65 ms 87.63 ms 10.26%
Gemma-4-31B-IT 463.89 ms 448.18 ms 3.39%
Nemotron-3-Nano-Omni 82.48 ms 74.08 ms 10.19%

Done tests: 11/11 focused W4A4 tests; 11/11 W4A8 and MoE tests; 238/238 supported broader FP4 cases;
CUDA memcheck reported zero errors on targeted cases.

Perplexity is very close, but not identical. Both tested models show slightly lower perplexity on this branch.

NVFP4 model Master Branch Change
Qwen3.6-35B-A3B 6.0428 6.0391 -0.061%
Qwen3.6-27B 7.1206 7.1114 -0.129%

Identical settings: WikiText-2 test data, first 16 chunks, context/batch/microbatch 4,096, Flash Attention enabled, full GPU offload. Each run evaluated 32,752 next-token predictions.

No perplexity regression appeared in this sample. Changing floating-point accumulation order can change rounding, so identical results are not guaranteed.

Similar results are achieved on DGX-Spark NVIDIA GB10, compute capability 12.1, VMM: yes, VRAM: 126758 MiB

n_ubatch Optimized (t/s) Master (t/s) Δ
512 2782.05 ± 15.08 2719.88 ± 14.04 +2.3%
1024 3092.86 ± 21.85 3030.23 ± 14.20 +2.1%
2048 3149.92 ± 31.97 3083.02 ± 22.37 +2.2%
4096 3150.17 ± 23.42 3080.44 ± 14.28 +2.3%

@kmorennv
kmorennv requested a review from a team as a code owner October 2, 2026 14:40
@ggml-gh-bot

ggml-gh-bot Bot commented Oct 2, 2026

Copy link
Copy Markdown

Hi @kmorennv, thanks for your contribution!

Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:

  • Multiple open PRs from a new contributor: We limit new contributors (those without a previously merged PR) to 1 open PR at a time. You currently have 2 open PRs.

Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below.

@kmorennv kmorennv changed the title Optimize accumulation in mmq for NVFP4 type GGML-Cuda: Optimize accumulation in mmq for NVFP4 type Oct 2, 2026
@github-actions github-actions Bot added ggml changes relating to the ggml tensor library for machine learning CUDA Related to the CUDA backend labels Oct 2, 2026
@kmorennv kmorennv changed the title GGML-Cuda: Optimize accumulation in mmq for NVFP4 type Cuda: Optimize accumulation in mmq for NVFP4 type Oct 2, 2026
@kmorennv kmorennv changed the title Cuda: Optimize accumulation in mmq for NVFP4 type CUDA: Optimize accumulation in mmq for NVFP4 type Oct 2, 2026

@ORippler ORippler left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks!

@ORippler
ORippler requested a review from am17an October 2, 2026 16:11
Comment thread ggml/src/ggml-cuda/mmq-vec-dot.cuh Outdated
Comment thread ggml/src/ggml-cuda/mmq-vec-dot.cuh Outdated
@ORippler
ORippler merged commit 8b2fbaf into ggml-org:master Oct 5, 2026
11 of 15 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants