cuda: bound the one-column MUL_MAT_ID mat-vec store by the expert row count - #295
Open
professorpalmer wants to merge 1 commit into
Open
professorpalmer wants to merge 1 commit into
professorpalmer wants to merge 1 commit into
Conversation
… count With ids and one token the expert slots are channels and stride_col_dst is nrows * n_used, so with small-K geometry (several rows per block) the last block of a slot stored its out-of-range rows into the next slot's first rows. Which write landed last decided the result, so the error was intermittent (PTQ1_0 m=70 k=2048: 67/300 runs failed on a 4070). The fused bias prefetch used the same bound. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
Fixes the intermittent
MUL_MAT_IDfailure reported on #221 (type_a=ptq1_0, n_mats=4, n_used=2, m=70, n=1, k=2048, error ~0.03 against 5e-4, on the base as well as the PR).Cause. For one-token
MUL_MAT_ID,ggml_cuda_mul_mat_vec_qputs the expert slots in channels and passesstride_col_dst = nb2 = nrows * n_used. The store guard inmul_mat_vec_qisrow0 + i < stride_col_dst, so when a block holds several rows andnrowsis not a multiple of that, the last block of slotsstores its out-of-range rows into the first rows of slots + 1. Two blocks write the same addresses and whichever lands last wins, which is why the failure is intermittent. At one column a block only holds several rows with the small-K geometry (should_use_small_k); PTQ1_0 takes it at k = 2048 (16 blocks per row).Fix. Bound the store (and the fused bias prefetch, which used the same guard) by the channel stride, which is the row count, in the one-column ids case. Other paths are unchanged.
The same guard is in ggml-org master (
mmvq.cu,row0 + i < stride_col_dst); other types reach it with short K, see below.Additional information
New
test-backend-opscases: every quantized type, one token, K = 2 blocks (forces small-K), m = 67,n_used2 and 4. The existing generic cases use m = 512, a multiple of every rows-per-block, so they could not hit the tail.RTX 4070 (sm_89), CUDA 13, Windows,
-b CUDA0:n_used=4)Full suites on the fixed build, default and
GGML_CUDA_BATCH_INVARIANT=1:MUL_MAT_ID1085/1085,MUL_MAT1516/1516,FLASH_ATTN_EXT2994/2994,GATED_DELTA_NET,GET_ROWSall pass.Requirements
🤖 Generated with Claude Code