Repository navigation
CUDA: Optimize accumulation in mmq for NVFP4 type - #29857
Merged
Merged
Conversation
|
Hi @kmorennv, thanks for your contribution! Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:
Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below. |
am17an
approved these changes
Oct 2, 2026
ORippler
reviewed
Oct 5, 2026
ORippler
approved these changes
Oct 5, 2026
1 task
This was referenced Oct 5, 2026
5 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
Simplified internal computation by accumulating results directly into the final output buffer instead of using a temporary intermediate step.
Additional information
The example to explain perf-gain use is Qwen3.6-35B-A3B-NVFP4, comparing to master
8d81559faagainst this branch:Settings: batch and microbatch 4,096, Flash Attention enabled, full GPU offload, RTX PRO 6000 Blackwell, CTK 13.4 Win64. Throughput used seven repetitions; nsys used three measured prompts after warmup.
The kernel improvement is visible across matched launch shapes. Both builds execute 860 NVFP4 MMQ launches per prompt. For example, one matching kernel/grid falls from 808.39 us to 751.55 us, a 7.03% reduction.
The smaller overall throughput gain also makes sense: NVFP4 MMQ accounts for approximately 31% of summed GPU kernel time in master. Reducing that component by 6.74% saves roughly 2.1% of total kernel time if everything else stays constant, consistent with the observed 2.65% prompt-throughput gain.
Nsight Compute confirms the effect across 12 matched NVFP4 launches from this model:
Across the other completed nsys captures, at the same batch size:
Done tests: 11/11 focused W4A4 tests; 11/11 W4A8 and MoE tests; 238/238 supported broader FP4 cases;
CUDA memcheck reported zero errors on targeted cases.
Perplexity is very close, but not identical. Both tested models show slightly lower perplexity on this branch.
Identical settings: WikiText-2 test data, first 16 chunks, context/batch/microbatch 4,096, Flash Attention enabled, full GPU offload. Each run evaluated 32,752 next-token predictions.
No perplexity regression appeared in this sample. Changing floating-point accumulation order can change rounding, so identical results are not guaranteed.
Similar results are achieved on DGX-Spark NVIDIA GB10, compute capability 12.1, VMM: yes, VRAM: 126758 MiB