Skip to content

metal: read FP16 and BF16 affine constants directly for FP32 inputs - #19

Merged
bri-prism merged 1 commit into
prismfrom
perf/qmv-mixed-scales
Sep 28, 2026
Merged

bri-prism merged 1 commit into
prismfrom
perf/qmv-mixed-scales

Conversation

@bri-prism

Copy link
Copy Markdown

With an FP32 activation and FP16 or BF16 scales and biases, quantized_matmul widens the constants on every call. This keeps them narrow on Metal: a new affine_qmv_fast_mixed kernel reads them in registers for the one-row path, and every other route widens them inside eval_gpu. Other backends and CPU streams are unchanged.

M5 Pro, Release, inference mode (train(false) asserted), plain greedy decode of a 2-bit signed-Hadamard target, 128 tokens, two alternating reps, tokens identical in every arm:

Arm Coding (ms/tok) Essay (ms/tok) Active memory
FP32 constant copies (current default) 34.61 / 34.64 34.17 / 34.15 8.644 GiB
FP16 constants read in the kernel 31.99 / 31.99 31.52 / 31.56 7.154 GiB

That is -7.6 percent per token and -1.49 GiB. DFlash2 does not regress (coding 13.13 vs 13.23 ms/tok, essay 69.6 vs 69.8).

Correctness: 193 of 193 cases are bitwise equal to widening first (FP16 and BF16 constants, 2, 3, 4 and 8 bits, M 1 to 64, several K and N including unaligned N, plus a CPU stream).

quantized_matmul used to cast an FP32 input's FP16 or BF16 scales and biases
to FP32 before every call: two extra kernels per projection, or a second FP32
copy of every constant kept resident. On Metal the op now keeps them narrow.
The one-row qmv_fast kernel reads them as they are and widens each value in
registers (a new affine_qmv_fast_mixed entry). Every other route widens them
inside eval_gpu first, as the op did. Outputs are bitwise identical to
widening first; other backends and CPU streams keep the old cast.
@bri-prism
bri-prism merged commit bbc151c into prism Sep 28, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant