Skip to content

sycl: XMX path for PQ2_0 and PTQ1_0 on 16-wide DPAS devices - #294

Merged
bri-prism merged 2 commits into
PrismML-Eng:prismfrom
kiljoy001:sycl-pq2-xe2-layout
Oct 2, 2026
Merged

bri-prism merged 2 commits into
PrismML-Eng:prismfrom
kiljoy001:sycl-pq2-xe2-layout

Conversation

@kiljoy001

@kiljoy001 kiljoy001 commented Sep 28, 2026 •

Copy link
Copy Markdown

Overview

Adds an XMX (DPAS) path for PQ2_0 and PTQ1_0 matrix multiplies on Intel GPUs with 16-wide DPAS: Xe-HPC (Data Center GPU Max), Xe2 (Battlemage, Lunar Lake), Xe3 (Panther Lake) and later. It covers every batch size: prompt processing, single-token decode, parallel decode and speculative verification. On those devices it replaces MMVQ and the dequantize-to-FP16 prompt path from #278 for these types.

The main idea: PQ2_0 packs value+1 in 2-bit fields, lowest first, which as a little-endian dword is already the DPAS 2-bit operand layout. The kernel feeds the codes to DPAS as signed 2-bit (s2) against int8 activations, converting them in registers with a borrow-free per-field subtract. Weights are never widened to 8 or 16 bits in memory.

How it works:

  • Weight layout. On first use, each PQ2_0 weight is rewritten in place into an SoA layout (same size): all 32-byte qs blocks first, then all fp16 scales. The qs row pitch is then nb*32 bytes, so one transposed 2D block load fetches 16 rows x 128 weights directly in the DPAS B-operand layout. The 34-byte AoS blocks are only 2-byte aligned and can't be loaded that way.
  • PTQ1_0. Base-3 has no DPAS form, so PTQ1_0 weights are expanded once into the same 2-bit layout. That is 34 bytes a block instead of 28. On devices that run this path the buffer type reserves the room for it (get_alloc_size), so VRAM holds only the expanded copy: about +25% for these weights, and the file on disk is unchanged.
  • Activations. They are quantized to int8 with one scale per 128 values, one PQ2_0 block. The four DPAS of a block then accumulate in int32 before a single float rescale, instead of one rescale per 32 values.
  • One kernel for all batch sizes. Tiles are 8/16/32 tokens x 32 rows. For small batches, up to 16 threads of a work-group split K for the same tile and reduce through SLM, so a mat-vec keeps enough loads in flight.

Scope and fallbacks:

  • On by default on any device whose int8 XMX DPAS is 16 wide, as the SYCL runtime reports it (matrix_combinations), not a list of device names. 8-wide XMX (Alchemist / Xe-HPG, Arrow Lake-H) and devices without XMX keep the existing paths; an 8-wide variant can follow separately once it can be tested on that hardware. GGML_SYCL_DISABLE_XMX=1 or GGML_SYCL_ENABLE_OPT=0 turns it off.
  • AOT builds (GGML_SYCL_DEVICE_ARCH) compile the kernels for every listed device, and they do not build for targets without 16-wide DPAS and 2D block loads. pq2_xmx.cpp is built only when every listed device is Xe-HPC, Xe2 or later, and compiles to stubs otherwise. Checked with ocloc: the kernels build for bmg-g21, bmg-g31, lnl-m, ptl-h, ptl-u, wcl, nvl-s, cri, pvc, xe-hpc, xe2-hpg, xe2-lpg, xe3-lpg, xe3p-xpc; the stubs build for acm-g10/g11/g12, ats-m150, dg1, tgllp, mtl-h, arl-h, pvc-vg.
  • Only plain MUL_MAT on 2D weights in regular (non-split, non-COMPUTE) buffers with K >= 256. Everything else (op offload, split buffers, MUL_MAT_ID, other devices) takes the existing paths unchanged.
  • Like the existing reorders, a tensor keeps its new layout once rewritten. The gate refuses MUL_MAT_ID expert slices and views. get_rows asserts if it ever sees a rewritten PQ2_0/PTQ1_0 tensor.

Files: pq2_xmx.cpp/.hpp (new), ggml-sycl.cpp (gate, alloc size, extras), common.hpp (layout flag, DPAS width), getrows.cpp (assert), CMakeLists.txt (AOT guard).

Results

Intel Arc Pro B50 (16 GB), oneAPI 2026.1, Level Zero, Linux xe driver, -ngl 99, Ternary Bonsai models. "#278" is the same build with GGML_SYCL_DISABLE_XMX=1, i.e. the current prism code path. Both columns were measured in the same session.

llama-bench, t/s:

model test #278 this PR change
1.7B PQ2_0 pp128 2574 6133 2.38x
pp512 5218 7645 1.47x
tg64 152.4 178.1 +17%
4B PQ2_0 pp128 1068 3068 2.87x
pp512 2298 3272 1.42x
tg64 77.4 99.8 +29%
27B PQ2_0 pp128 148 346 2.33x
pp512 290 370 1.28x
tg64 12.39 16.53 +33%
27B PTQ1_0 pp128 127 348 2.74x
pp512 267 370 1.39x
tg64 12.88 16.54 +28%

Parallel decode (llama-batched-bench -npp 128 -ntg 64, aggregate t/s):

sequences 1.7B PQ2_0 4B PQ2_0 27B PQ2_0 27B PTQ1_0
1 151 -> 171 76 -> 97 12.2 -> 16.1 13.1 -> 16.1
2 250 -> 334 125 -> 186 18.9 -> 29.7 20.2 -> 29.7
4 380 -> 613 185 -> 355 25.2 -> 48.0 26.3 -> 47.9
8 478 -> 1010 224 -> 618 29.0 -> 68.6 31.0 -> 68.0
16 322 -> 1446 138 -> 900 20.2 -> 85.7 17.2 -> 85.8
32 588 -> 1769 262 -> 1184 32.5 -> 98.4 29.0 -> 98.4

On #278, batches of 9+ tokens leave MMVQ for the dequantize-to-FP16 path, which is why throughput drops at 16 sequences there.

Speculative decoding (27B PQ2_0 target, Qwen3.5-0.8B Q8_0 draft, --spec-draft-n-max 1, greedy, 256 tokens), t/s:

#278 this PR
prose, target only 13.1 17.6
prose, with draft 12.1 20.5
code, target only 12.7 17.5
code, with draft 12.8 21.9

With #278, drafting does not pay. With this PR, verification is cheap enough that n_max=1 gains 17-25% over the target alone. Longer drafts lose on acceptance (15-33% at n_max 2-3 with this draft model).

Where the decode gain comes from: at the 27B FFN shape (17408x5120, one token), the PQ2_0 mat-vec goes from 199 us (MMVQ, about 53% of the B50's 224 GB/s) to 88 us. End to end, 27B decode now reads about 123 GB/s. Most of the remainder is the ~2300 non-matmul kernels per token.

Correctness

  • test-backend-ops -b SYCL0: MUL_MAT 1329/1329, MUL_MAT_ID 1035/1035, GET_ROWS 119/119 (all types). Full suite, run per op: 14703/14759. The 56 failures (CONV_2D, ROLL, FLASH_ATTN_EXT, LIGHTNING_INDEXER) are identical with this path disabled, so they predate this PR. SET_ROWS aborts on current prism regardless of this PR (set_rows.cpp:549: Unsupported tensor type!), so it was run separately.
  • Perplexity (-c 512, 5 chunks), sycl: 2-3x prompt processing speedup by dequantizing PTQ1_0 and PQ2_0 to FP16 #278 -> this PR:
model setting #278 this PR
1.7B PQ2_0 b1 8.4089 8.4335
1.7B PQ2_0 b512 8.4233 8.4149
4B PQ2_0 b1 7.0584 7.0565
4B PQ2_0 b512 7.0545 7.0506
1.7B PQ2_0 -ngl 0 (op offload) 8.4225 8.4225
4B PQ2_0 -ngl 20 (partial) 7.0541 7.0535
27B PTQ1_0 b512 4.2421 4.2399

Additional information

  • This introduces ESIMD/XMX code to the SYCL backend, a new pattern here. I'm happy to adjust structure or naming to fit how the maintainers want it organized.
  • Tuning (tile sizes, K-split thresholds) was done on the B50 only. Other 16-wide DPAS devices are enabled by capability but untested on hardware. A UHD 770 (no XMX) in the same machine stays on the existing paths and passes MUL_MAT for PQ2_0/PTQ1_0.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES - the code and this description were written by Claude (Anthropic) under my direction; I tested it on my own hardware (Intel Arc Pro B50) and reviewed the changes.

…el Xe2 GPUs (Battlemage, Lunar Lake). It covers every batch size: prompt processing, single-token decode, parallel decode and speculative verification. On Xe2 it replaces MMVQ and the dequantize-to-FP16 prompt path from ggml-org#278 for these types.

The main idea: PQ2_0 packs value+1 in 2-bit fields, lowest first, which as a little-endian dword is already the DPAS 2-bit operand layout. The kernel feeds the codes to DPAS as signed 2-bit (s2) against int8 activations, converting them in registers with a borrow-free per-field subtract. Weights are never widened to 8 or 16 bits in memory.

How it works:
- **Weight layout.** On first use, each PQ2_0 weight is rewritten in place into an SoA layout (same size): all 32-byte qs blocks first, then all fp16 scales. The qs row pitch is then nb*32 bytes, so one transposed 2D block load fetches 16 rows x 128 weights directly in the DPAS B-operand layout. The 34-byte AoS blocks are only 2-byte aligned and can't be loaded that way.
- **PTQ1_0.** Base-3 has no DPAS form, so PTQ1_0 weights are expanded once into the same 2-bit layout. That is 34 bytes a block instead of 28. On Xe2 devices the buffer type reserves the room for it (get_alloc_size), so VRAM holds only the expanded copy: about +25% for these weights, and the file on disk is unchanged.
- **Activations.** They are quantized to int8 with one scale per 128 values, one PQ2_0 block. The four DPAS of a block then accumulate in int32 before a single float rescale, instead of one rescale per 32 values.
- **One kernel for all batch sizes.** Tiles are 8/16/32 tokens x 32 rows. For small batches, up to 16 threads of a work-group split K for the same tile and reduce through SLM, so a mat-vec keeps enough loads in flight.

Scope and fallbacks:
- On by default on BMG-G21/G31 and LNL-M. `GGML_SYCL_DISABLE_XMX=1` or `GGML_SYCL_ENABLE_OPT=0` turns it off.
- Only plain MUL_MAT on 2D weights in regular (non-split, non-COMPUTE) buffers with K >= 256. Everything else (op offload, split buffers, MUL_MAT_ID, other devices) takes the existing paths unchanged.
- Like the existing reorders, a tensor keeps its new layout once rewritten. The gate refuses MUL_MAT_ID expert slices and views. `get_rows` asserts if it ever sees a rewritten PQ2_0/PTQ1_0 tensor.

Files: `pq2_xe2.cpp/.hpp` (new), `ggml-sycl.cpp` (gate, alloc size, extras), `common.hpp` (layout flag), `getrows.cpp` (assert).

Intel Arc Pro B50 (16 GB), oneAPI 2026.1, Level Zero, Linux `xe` driver, `-ngl 99`, Ternary Bonsai models. "ggml-org#278" is the same build with `GGML_SYCL_DISABLE_XMX=1`, i.e. the current prism code path. Both columns were measured in the same session.

`llama-bench`, t/s:

| model | test | ggml-org#278 | this PR | change |
|---|---|---|---|---|
| 1.7B PQ2_0 | pp128 | 2574 | 6133 | 2.38x |
| | pp512 | 5218 | 7645 | 1.47x |
| | tg64 | 152.4 | 178.1 | +17% |
| 4B PQ2_0 | pp128 | 1068 | 3068 | 2.87x |
| | pp512 | 2298 | 3272 | 1.42x |
| | tg64 | 77.4 | 99.8 | +29% |
| 27B PQ2_0 | pp128 | 148 | 346 | 2.33x |
| | pp512 | 290 | 370 | 1.28x |
| | tg64 | 12.39 | 16.53 | +33% |
| 27B PTQ1_0 | pp128 | 127 | 348 | 2.74x |
| | pp512 | 267 | 370 | 1.39x |
| | tg64 | 12.88 | 16.54 | +28% |

Parallel decode (`llama-batched-bench -npp 128 -ntg 64`, aggregate t/s):

| sequences | 1.7B PQ2_0 | 4B PQ2_0 | 27B PQ2_0 | 27B PTQ1_0 |
|---|---|---|---|---|
| 1 | 151 -> 171 | 76 -> 97 | 12.2 -> 16.1 | 13.1 -> 16.1 |
| 2 | 250 -> 334 | 125 -> 186 | 18.9 -> 29.7 | 20.2 -> 29.7 |
| 4 | 380 -> 613 | 185 -> 355 | 25.2 -> 48.0 | 26.3 -> 47.9 |
| 8 | 478 -> 1010 | 224 -> 618 | 29.0 -> 68.6 | 31.0 -> 68.0 |
| 16 | 322 -> 1446 | 138 -> 900 | 20.2 -> 85.7 | 17.2 -> 85.8 |
| 32 | 588 -> 1769 | 262 -> 1184 | 32.5 -> 98.4 | 29.0 -> 98.4 |

On ggml-org#278, batches of 9+ tokens leave MMVQ for the dequantize-to-FP16 path, which is why throughput drops at 16 sequences there.

Speculative decoding (27B PQ2_0 target, Qwen3.5-0.8B Q8_0 draft, `--spec-draft-n-max 1`, greedy, 256 tokens), t/s:

| | ggml-org#278 | this PR |
|---|---|---|
| prose, target only | 13.1 | 17.6 |
| prose, with draft | 12.1 | 20.5 |
| code, target only | 12.7 | 17.5 |
| code, with draft | 12.8 | 21.9 |

With ggml-org#278, drafting does not pay. With this PR, verification is cheap enough that n_max=1 gains 17-25% over the target alone. Longer drafts lose on acceptance (15-33% at n_max 2-3 with this draft model).

Where the decode gain comes from: at the 27B FFN shape (17408x5120, one token), the PQ2_0 mat-vec goes from 199 us (MMVQ, about 53% of the B50's 224 GB/s) to 88 us. End to end, 27B decode now reads about 123 GB/s. Most of the remainder is the ~2300 non-matmul kernels per token.

- `test-backend-ops -b SYCL0`: MUL_MAT 1329/1329, MUL_MAT_ID 1035/1035, GET_ROWS 119/119 (all types). Full suite, run per op: 14703/14759. The 56 failures (CONV_2D, ROLL, FLASH_ATTN_EXT, LIGHTNING_INDEXER) are identical with this path disabled, so they predate this PR. SET_ROWS aborts on current prism regardless of this PR (`set_rows.cpp:549: Unsupported tensor type!`), so it was run separately.
- Perplexity (`-c 512`, 5 chunks), ggml-org#278 -> this PR:

| model | setting | ggml-org#278 | this PR |
|---|---|---|---|
| 1.7B PQ2_0 | b1 | 8.4089 | 8.4335 |
| 1.7B PQ2_0 | b512 | 8.4233 | 8.4149 |
| 4B PQ2_0 | b1 | 7.0584 | 7.0565 |
| 4B PQ2_0 | b512 | 7.0545 | 7.0506 |
| 1.7B PQ2_0 | -ngl 0 (op offload) | 8.4225 | 8.4225 |
| 4B PQ2_0 | -ngl 20 (partial) | 7.0541 | 7.0535 |
| 27B PTQ1_0 | b512 | 4.2421 | 4.2399 |

- KL divergence against a CPU-only reference (1.7B, `-dev none --no-op-offload`), ggml-org#278 vs this PR:
  - b1: 0.00122 vs 0.00125.
  - b512: 0.00106 vs 0.00113.
  - Same top token: 97.4-97.8% vs 97.7-97.9%.
  - The differences are within the error bars. The per-128 activation scale costs nothing measurable.

- This introduces ESIMD/XMX code to the SYCL backend, a new pattern here. I'm happy to adjust structure or naming to fit how the maintainers want it organized.
- Tuning (tile sizes, K-split thresholds) was done on the B50 only. BMG-G31 and LNL-M are gated in by architecture but untested.

- I have read and agree with the [contributing guidelines](https://github.com/ggml-org/llama.cpp/blob/master/CONTRIBUTING.md)
- AI usage disclosure: YES - the code and this description were written by Claude (Anthropic) under my direction; I tested it on my own hardware (Intel Arc Pro B50) and reviewed the changes.
@akapug

akapug commented Oct 1, 2026

Copy link
Copy Markdown

B70 numbers for this PR, since the description has a B50 only. They are from one card, one model and one session, with no repeat window.

Setup: Arc Pro B70 (BMG-G31), Linux xe, Level Zero 1.15.38308, oneAPI 2026.1.1, prism 88c4bc6 with this PR's head 09385c4 merged. Model: Ternary-Bonsai-2-27B PTQ1_0, -ngl 99 -fa on. The "#278" column is the same build with GGML_SYCL_DISABLE_XMX=1.

Correctness.

  • test-backend-ops -o MUL_MAT: 1401/1401 with this PR. With XMX disabled it is 1400/1401; the one failure is MUL_MAT(type_a=pq2_0,type_b=f32,m=16,n=1,k=256), ERR 0.000577 > 0.0005. So it is on the existing path, not this PR.
  • Greedy 128-token chat output is identical between the two.

llama-bench, t/s (-r 5):

test #278 this PR
pp128 326.4 784.8
pp512 691.7 924.5
tg64 26.99 32.70

Parallel decode (llama-batched-bench -npp 128 -ntg 64, aggregate S_TG t/s):

sequences #278 this PR
2 46.2 58.2
4 62.9 102.4
8 76.3 156.3

The 1-sequence row read 22.2 with this PR, below its own tg64. I have not looked into why.

One data point that may matter for the decode path. We serve this model on a B70 with our own PTQ1_0 mat-vec. It keeps the 28-byte base-3 blocks (no expansion) and runs one row per lane. I timed it against this PR's kernel and MMVQ at the 27B FFN shape (17408x5120). Each graph uses 8 distinct weights, so every weight comes from DRAM, not cache. Median µs per mat-vec:

tokens #278 (MMVQ) this PR base-3 mat-vec
1 88.8 64.0 40.1
2 100.0 63.0 39.6
4 149.8 65.0 42.4
8 241.6 66.8 56.8

That is about 80 % of the B70's 608 GB/s at 1-2 tokens for the base-3 kernel, against about 61 % for this PR's kernel, counting its expanded 2-bit bytes. This PR's time stays flat as tokens grow, so it should overtake somewhere above 8. It is also well ahead on short prompts (pp128).

Would you want the base-3 mat-vec as the path for up to ~8 tokens, beside this PR's DPAS path? The catch is that both rewrite the weight in place, in different layouts. One tensor can't serve both without a second copy or converting on the fly. I'm happy to open a PR against this branch, or against prism after this lands, whichever suits you.

-- Claude Opus 5.5, working in helm for @akapug

@bri-prism

Copy link
Copy Markdown
Collaborator

Tested on an Arc B390 (Panther Lake, Xe3). As submitted it's a no-op there because the XMX gate only admits BMG and LNL device ids. With PTL added to the gate (intel_gpu_ptl_h and ptl_u) it passes test-backend-ops and gives a large prefill and batched-decode gain with KLD at noise level. Could you add PTL to the gate? Happy to retest.

…ice list

The path was limited to BMG-G21/G31 and LNL-M by name. It now runs on any XMX device whose int8 DPAS is 16 wide,
as the runtime reports it through matrix_combinations: Xe-HPC, Xe2, Xe3 and later. 8-wide XMX (Xe-HPG, Arrow
Lake-H) and devices without XMX keep the existing paths.

AOT builds compile the kernels for every GGML_SYCL_DEVICE_ARCH device, and they do not build for targets without
16-wide DPAS and 2D block loads (ocloc crashes on acm-g10, and fails on tgllp, dg1, mtl and pvc-vg). pq2_xmx.cpp
is now built only when every AOT device is Xe-HPC, Xe2 or later; otherwise it compiles to stubs.

Renames pq2_xe2 to pq2_xmx to match.
@kiljoy001 kiljoy001 changed the title sycl: XMX path for PQ2_0 and PTQ1_0 on Xe2 sycl: XMX path for PQ2_0 and PTQ1_0 on 16-wide DPAS devices Oct 1, 2026
@bri-prism

Copy link
Copy Markdown
Collaborator

Retested 3d6ca77 on the Arc B390 (Panther Lake, Xe3) without any local patch: the DPAS-16 gate now selects the XMX path there by itself and gives the same results as my patched build. test-backend-ops SYCL0 MUL_MAT 1401/1401, MUL_MAT_ID 1035/1035, GET_ROWS 119/119; KLD against the FP16 path 1e-4 to 1e-3. On the 27B, decode is +13% (PTQ1_0) and +21% (PQ2_0), and batched decode at 16 streams goes from 13.6 to 65.1 t/s. Thanks for the quick fix, this looks good from the Intel side.

@bri-prism
bri-prism merged commit a14c7de into PrismML-Eng:prism Oct 2, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants