Conversation
On Intel GPUs the MMVQ path for both ternary formats now runs an ESIMD kernel that gives each work-item one weight row, spreads the row's blocks over 16 lanes, and accumulates q8_1 dot products with dp4a. The kernels are compiled under the Intel compiler only and follow GGML_SYCL_ENABLE_ESIMD, so setting it to 0 restores the sub-group kernels for comparison. A one-row GET_ROWS with a single index becomes a vectorized copy, which is the shape the recurrent-state layers issue every token. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DSrpF2XYi9dav1e3hSAkaS Claude-Session: https://claude.ai/code/session_014smgAQnqKxMyRYQbXnvBgs
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
On Intel GPUs, single-token decode of PTQ1_0 and PQ2_0 runs through MMVQ, where the sub-group kernels from PrismML-Eng#235 set the decode speed. This PR adds explicit-SIMD (ESIMD) MMVQ kernels for both formats: each work-item takes one weight row, spreads the row's blocks over 16 lanes, and accumulates q8_1 dot products with dp4a, four rows per work-group. They compile only under the Intel compiler and follow a new
GGML_SYCL_ENABLE_ESIMDswitch, default 1; setting it to 0 restores the sub-group kernels.It also turns a one-row, single-index F32 GET_ROWS into a vectorized copy. The recurrent-state layers of the qwen35 graph issue that shape once per layer per token.
Prompt processing is unchanged: batches above the MMVQ width take the FP16 GEMM from PrismML-Eng#278.
Additional information
Intel Arc 140V (Lunar Lake iGPU), Windows, oneAPI 2025.3, driver 32.0.101.9030, Release. Base
prism@ 88c4bc6, Ternary-Bonsai-2-27B.llama-bench -ngl 99 -fa 1 -p 512 -n 128 -r 1, four rounds in separate processes 30 s apart with the arm order flipped each round; median (range) in t/s:With
GGML_SYCL_ENABLE_ESIMD=0, this PR decodes PTQ1_0 at 6.07 and 6.08 t/s, against 6.99 with it on. This laptop lowers its memory clock under sustained load, which accounts for the spread in prompt speed.Correctness:
test-backend-ops -o MUL_MAT: PQ2_0 216/216 and PTQ1_0 167/167 on base and this PR, with oneDNN, withGGML_SYCL_ENABLE_DNN=0, and withGGML_SYCL_ENABLE_ESIMD=0. 48 cases per type have n = 1 at the model's row lengths and take the new kernels.-o GET_ROWS -p type=f32: 11/11.GGML_SYCL_ENABLE_ESIMD=0both formats match base exactly, so the GET_ROWS copy changes nothing.llama-perplexity --kl-divergence, WikiText-2,-c 2048 -b 512, 8 chunks, against base: mean KLD 0.000000, max KLD 0.000056, same top 100% for both formats. These batches take the GEMM path, so this mainly confirms prompt processing is untouched.Only tested on Xe2.
Requirements