Conversation
…GEMM On Intel GPUs the MMVQ path for both ternary formats now runs an ESIMD kernel that gives each work-item one weight row, spreads the row's blocks over 16 lanes, and accumulates q8_1 dot products with dp4a. The kernels are compiled under the Intel compiler only and follow GGML_SYCL_ENABLE_ESIMD, so setting it to 0 restores the sub-group kernels for comparison. Prompt batches of both formats dequantize to FP16 before the GEMM, and a one-row GET_ROWS with a single index becomes a vectorized copy, which is the shape the recurrent-state layers issue every token. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DSrpF2XYi9dav1e3hSAkaS
Owner
Author
|
Superseded by #8, which carries this PR's decode kernels and GET_ROWS copy rebased onto current prism head. The FP16 prompt change from this PR merged upstream separately as PrismML-Eng#278. Written by Claude Code on Daniel's instruction. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR adds explicit-SIMD decode kernels for PTQ1_0 and PQ2_0 to the SYCL backend, runs prompt batches of both formats through an FP16-input GEMM, and turns the one-row recurrent-state GET_ROWS into a vectorized copy. The kernels follow GGML_SYCL_ENABLE_ESIMD; setting it to 0 restores the sub-group kernels.
Arc 140V at default settings, against prism head, in tokens per second: PQ2_0 decode 6.8 → 8.8, prompt 55 → 88–118; PTQ1_0 decode 5.9 → 7.0, prompt 57 → 147. Over eight 2,048-token WikiText chunks, mean KL divergence from prism head is 0.000001 with 99.96% same top token.