Skip to content

vulkan: integer-dot mat-vec kernel for PTQ1_0 (8x faster decode) and PQ2_0 support - #188

Closed
alhnesn wants to merge 4 commits into
PrismML-Eng:prismfrom
alhnesn:vulkan-ptq1_0-mmvq-pq2_0
Closed

alhnesn wants to merge 4 commits into
PrismML-Eng:prismfrom
alhnesn:vulkan-ptq1_0-mmvq-pq2_0

Conversation

@alhnesn

@alhnesn alhnesn commented Sep 18, 2026 •

Copy link
Copy Markdown

Summary

On Vulkan, PTQ1_0 decode currently goes through the generic mul_mat_vec.comp path, which calls ptq1_0_trit() per element (a base-3 loop with a branch per trit). On an RX 6750 XT that shader runs at ~270 GFLOPS and Ternary-Bonsai-2-27B decodes at 4.7 tok/s, while the same card streams Q4_0 weights at ~330 GB/s. PQ2_0 has no Vulkan support at all and falls back to the CPU.

This PR adds:

  • mul_mat_vecq_ptq1_0.comp: a q8_1 (mmvq) shader for PTQ1_0. Two lanes share one 128-element block (lane 0: qs dwords 0..2, lane 1: dwords 3..5 + qh), 4 rows per workgroup, identical instruction stream on both lanes. Trits are decoded with the multiply-by-3 recurrence on two bytes at a time in the 16-bit halves of a dword (same idea as vec_dot_ptq1_0_q8_1 in the CUDA backend). The level word is fed to dotPacked4x8AccSatEXT without extracting the trit bits: the activation words are pre-masked so the residue bytes multiply by zero. The ternary -1 offset is folded into the q8_1 block sums.
  • PQ2_0 in the Vulkan backend: block and packed16 types, dequant_pq2_0.comp, mul_mm loader, float and q8_1 mat-vec paths (modeled on the existing Q2_0 code with the 128-element group), get_rows, and registration in vulkan-shaders-gen.cpp / ggml-vulkan.cpp. Like PTQ1_0 it has no coopmat2 decoder and is skipped for cm2 generation.
  • Tests: test-backend-ops MUL_MAT / MUL_MAT_ID cases for both types at the Bonsai shapes (k = 1024, 5120, 6144, 17408), odd row counts (m = 67 / 70), batched B, n = 1..8, and perf-mode cases at Bonsai shapes.

Results (RX 6750 XT 12 GB, RDNA2, AMD proprietary driver, Windows 11, Ternary-Bonsai-2-27B)

before after
PTQ1_0 decode, -fa off 4.7 tok/s 40.8 tok/s
PTQ1_0 decode, -fa on 4.7 tok/s 39.0 tok/s
PTQ1_0 decode at 8k context – 34.1 tok/s
PQ2_0 decode CPU fallback 38.5 tok/s
PQ2_0 prefill (pp512) CPU fallback 200 t/s

The PTQ1_0 kernel streams weights at ~330 GB/s (card peak 432 GB/s); the remaining per-token time is the model's other ops.

Batched decode (second commit)

The first version re-ran the trit decode for every column of B, so several sequences per step (llama-server with -np N) scaled poorly. The second commit adds a multi-column path: columns are handled in passes of 3, and per block the lane decodes all rows once, level by level, and dots each packed level with the activations of every column of the pass. The single-column pipeline keeps the original loop unchanged. llama-batched-bench -npp 256 -ntg 128 -fa on -ctk q8_0 -ctv q8_0, aggregate decode:

parallel sequences first commit second commit per sequence
1 36.6 tok/s 38.1 tok/s 38.1
2 39.7 tok/s 60.8 tok/s 30.4
3 – 75.3 tok/s 25.1
4 55.2 tok/s 73.3 tok/s 18.3
8 66.3 tok/s 86.4 tok/s 10.8

Kernel time for m = 17408, k = 5120: 45 µs at n = 1, 57 at n = 2, 120 at n = 4, 215 at n = 8 (before: n = 8 took about 5× n = 1). Two things worth knowing if you tune this further on other hardware: the driver hoists the loads of everything in one block body, so all passes had to be separate loops (one loop with all columns spilled at n = 8), and it does not unroll a loop that contains the block loop, which is why the passes are guarded calls. A 4-column pass was slower than 3 + 1 on this card for reasons I could not pin down (not registers).

Verification

  • test-backend-ops test -o MUL_MAT,MUL_MAT_ID,GET_ROWS against the CPU backend: all cases pass for both types, on the default paths and with GGML_VK_FORCE_MMVQ=1, with the subgroup-size workgroup and (by forcing DMMV_WG_SIZE_LARGE locally) the large-workgroup variant.
  • Full MUL_MAT / MUL_MAT_ID / GET_ROWS suite across all types passes (default and forced mmvq), so the shared-file changes (types.glsl, dequant_funcs.glsl, mul_mm_funcs.glsl, mul_mat_vecq*.glsl) don't regress other formats.
  • Perplexity through the mat-vec path (-ub 8): 6.3916 with the new shader vs 6.3906 with the original float path; greedy output token-identical to the stock prism-b10685 Vulkan build over 80 tokens. After the second commit the same -ub 8 run (now through the multi-column path) gives 6.2321 vs 6.2294 for the first commit on a 12-chunk subset, the difference being fp32 summation order.
  • Second commit: 271/271 test-backend-ops cases for PTQ1_0 + PQ2_0 with forced mmvq (n = 1..8 including the row tail), 138/138 on the default path, and the pipeline statistics extension confirms no scratch use at any column count (72 VGPRs at n = 1, at most 127 at n = 8).
  • PQ2_0 and PTQ1_0 give identical perplexity and greedy text, as expected for the same ternary weights.

Caveats

  • Only tested on one GPU (gfx1031, subgroup size 32, int8 dot support). The shader is written for any BLOCK_SIZE and both reduction modes, but NVIDIA / Intel / RADV runs would be very welcome.
  • ggml_vk_should_use_mmvq is unchanged: the mmvq path is taken by the existing vendor heuristics (on AMD for k >= 2048), otherwise the float path is used as before.
  • Variants I measured that were slower on RDNA2 and are not included: shared-memory staging of the weight slab, 8 rows per workgroup, splitting the workgroup across rows, next-block prefetch, and a 4-lanes-per-block layout with per-lane unit tables.

AI use

This kernel and the PQ2_0 port were developed using Claude Fable 5.1 in Claude Code, working interactively with the author on the target machine.

PTQ1_0 decode on Vulkan went through the generic mul_mat_vec shader, which
decodes every trit with a per-element base-3 loop; on an RX 6750 XT that ran
at ~270 GFLOPS and limited Ternary-Bonsai-2-27B to 4.7 tok/s. This adds a
dedicated q8_1 (mmvq) shader for PTQ1_0 and wires PQ2_0 into the Vulkan
backend, which had no support for it at all.

mul_mat_vecq_ptq1_0.comp:
- two lanes per 128-element block, 4 rows per workgroup, no divergence
- trits decoded with the multiply-by-3 recurrence on two bytes at a time in
  the 16-bit halves of a dword (as the CUDA vec_dot_ptq1_0_q8_1 does)
- the level word is fed to dotPacked4x8AccSat as is: activations are
  pre-masked so the residue bytes are multiplied by zero, which removes the
  per-level extraction ops
- the ternary -1 offset is folded into the q8_1 block sums

PQ2_0: block/packed16 types, dequant shader, mul_mm loader, float and q8_1
mat-vec paths (modeled on Q2_0 with the 128-element group), get_rows.

test-backend-ops: MUL_MAT / MUL_MAT_ID cases for both types at the Bonsai
shapes (k = 1024..17408), odd row counts, batched B and n = 1..8, plus perf
cases at Bonsai shapes.

Measured on RX 6750 XT (RDNA2, AMD proprietary driver, Windows 11),
Ternary-Bonsai-2-27B: PTQ1_0 decode 4.7 -> 40.8 tok/s (39 with -fa on),
PQ2_0 CPU-only -> 38.5 tok/s, prefill 200 t/s. Perplexity through the new
path 6.3916 vs 6.3906 with the original shader; greedy output identical to
the stock build.
The PTQ1_0 q8_1 mat-vec re-ran the whole trit decode for every column of B,
so batched decode (llama-server with several slots, NUM_COLS = 2..8) scaled
poorly: on an RX 6750 XT 8 parallel sequences gave 66 tok/s aggregate, barely
more than 4.

Columns are now handled in passes of CG = 3 (a new specialization constant).
Per block the lane decodes all NUM_ROWS rows level by level once and dots each
packed level with the activations of every column of the pass, so an extra
column costs three loads and three dot products per row and level. Integer
accumulators are folded into the fp32 result whenever the level moves on to a
new q8_1 sub-block scale. Passes are separate loops over the blocks, written
as guarded calls, because the driver hoists the loads of everything in one
block body (which spilled at 8 columns) and does not unroll a loop that
contains the block loop.

The single-column pipeline keeps the original loop (runtime-bounded row loop,
72 VGPRs), so single-sequence decode is unchanged.

RX 6750 XT, Ternary Bonsai 2 27B, llama-batched-bench pp256/tg128, fa on,
q8_0 KV, aggregate decode tok/s (before -> after):
  1 seq  36.6 -> 38.1     4 seq  55.2 -> 73.3
  2 seq  39.7 -> 60.8     8 seq  66.3 -> 86.4
llama-bench tg64: 40.6 (fa off) / 38.7 (fa on), same as before.
Kernel time m=17408 k=5120: n=1 45 us, n=2 57, n=4 120, n=8 215.

test-backend-ops: 271/271 (ptq1_0 + pq2_0, forced mmvq, n = 1..8 incl. row
tail), 138/138 default path; perplexity through the 8-column path 6.2321 vs
6.2294 before (same text, fp32 summation order).

Also adds perf cases with n = 2, 4, 8 for PTQ1_0 / PQ2_0 / Q4_0.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

The new PTQ1_0 and PQ2_0 MUL_MAT_ID integer-dot pipelines are currently unreachable from the pipeline selector.

Get a fresh assessment by requesting another Copilot review.

Pull request overview

Adds optimized Vulkan PTQ1_0 integer-dot decoding and PQ2_0 backend support.

Changes:

  • Adds PTQ1_0 single- and multi-column integer-dot shaders.
  • Adds PQ2_0 dequantization, matmul, mat-vec, and row-access support.
  • Expands correctness and performance tests.
File summaries
File Description
tests/test-backend-ops.cpp Adds PTQ1_0/PQ2_0 tests and benchmarks.
ggml/src/ggml-vulkan/vulkan-shaders/vulkan-shaders-gen.cpp Generates and exports new shader variants.
ggml/src/ggml-vulkan/vulkan-shaders/types.glsl Defines PQ2_0 shader layouts.
ggml/src/ggml-vulkan/vulkan-shaders/mul_mm_funcs.glsl Loads PQ2_0 matrix tiles.
ggml/src/ggml-vulkan/vulkan-shaders/mul_mat_vecq.comp Enables PQ2_0 integer-dot dispatch.
ggml/src/ggml-vulkan/vulkan-shaders/mul_mat_vecq_ptq1_0.comp Implements optimized PTQ1_0 mat-vec.
ggml/src/ggml-vulkan/vulkan-shaders/mul_mat_vecq_funcs.glsl Implements PQ2_0 integer-dot helpers.
ggml/src/ggml-vulkan/vulkan-shaders/dequant_pq2_0.comp Adds standalone PQ2_0 dequantization.
ggml/src/ggml-vulkan/vulkan-shaders/dequant_funcs.glsl Adds shared PQ2_0 decoding helpers.
ggml/src/ggml-vulkan/ggml-vulkan.cpp Registers and selects PQ2_0/PTQ1_0 pipelines.
Review details
  • Files reviewed: 10/10 changed files
  • Comments generated: 4
  • Review effort level: Balanced

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread ggml/src/ggml-vulkan/ggml-vulkan.cpp
Comment thread ggml/src/ggml-vulkan/vulkan-shaders/mul_mat_vecq_ptq1_0.comp Outdated
Comment thread ggml/src/ggml-vulkan/vulkan-shaders/vulkan-shaders-gen.cpp
Comment thread tests/test-backend-ops.cpp Outdated
@drien

drien commented Sep 18, 2026

Copy link
Copy Markdown

PQ2_0 works great on Linux/RADV with an RDNA2 card. Tested at e80f16d (both commits):

  • GPU: RX 6900 XT (gfx1030 / RADV NAVI21), Mesa 26.0.3, kernel 7.0.0, glslc shaderc 2026.1 (int dot: 1, warp size: 32, no coopmat)
  • test-backend-ops test -b Vulkan0 -o MUL_MAT,MUL_MAT_ID,GET_ROWS: 2279/2279 pass on the default path; 2152/2152 with GGML_VK_FORCE_MMVQ=1. No PQ2_0 or PTQ1_0 failures in either.
  • llama-bench -m Ternary-Bonsai-2-27B-PQ2_0.gguf -ngl 99: pp512 = 308 t/s, tg64 = 44.1 t/s. Before this PR the same model fell back to the CPU (55 MiB on device, 902 graph splits) at ~3 t/s.
  • llama-server with -c 65536 (n_parallel 4): 6.5 GiB of weights on Vulkan0, 2 graph splits, coherent output at temp 0.

@MarkPronkin

Copy link
Copy Markdown

Tested this on my RX 580, went from 1 token a second to roughly 6.

…ments

- ggml_vk_get_dequantize_mul_mat_vec_id: add PTQ1_0 and PQ2_0 to the q8_1
  whitelist. The id q8_1 pipelines were created but never selected, so
  MUL_MAT_ID always fell back to the float path.
- test-backend-ops: add n = 4 to the PTQ1_0/PQ2_0 cases (a three-column pass
  followed by the single-column tail).
- Shorten the shader comments and the coopmat2 note in vulkan-shaders-gen.cpp.

test-backend-ops MUL_MAT/MUL_MAT_ID/GET_ROWS for ptq1_0 + pq2_0: 284/284 on
the default path and 284/284 with GGML_VK_FORCE_MMVQ=1 (RX 6750 XT); the
mul_mat_vec_id_ptq1_0_q8_1_f32 pipeline is now compiled and used.
@rofl0r

rofl0r commented Sep 19, 2026

Copy link
Copy Markdown

tested the PTQ1_0 on intel xe igpu, works great - inference 11 t/s which is more than 3x of what i get with qwen 3.8 27b q4_k_m. however, speed degrades quickly to ~6 t/s at 30K context, and prefill speed it pretty slow - starts at 80t/s empty and degrades to 30 t/s at 30K context.

@bri-prism bri-prism left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agent review: posted by the maintainer's coding agent at their request.

No additional confirmed finding in this source pass. The current patch addresses the earlier review's missing PTQ1_0/PQ2_0 ID q8_1 whitelist entries and expands the column-count test coverage.

Vulkan compilation and numerical execution were not performed here. Retain forced-MMVQ MUL_MAT_ID coverage and all column counts 1-8, especially the three-column group plus single-column tail, as the merge gate. No independent performance claim is made by this review.

Reviewed commit: 330e828b482b208d66f83393146aaae2d68553bd.

@renovys

renovys commented Sep 19, 2026

Copy link
Copy Markdown

Tested this PR on an AMD BC-250 (gfx1013, RADV/Mesa, Vulkan UMA ~13.6 GB; device
caps: fp16 yes, int dot no, matrix cores none, warp 32) with
Ternary-Bonsai-2-27B in PTQ_1.75 (PTQ1_0) and PQ2_0.

Builds: prism base 9a9394a89 + the three #188 commits (f48175125 head), and
the same plus #187 (b1b211a3d head). llama-bench -ngl 99 -fa 1 -p 512 -n 128 -r 3:

build PTQ1_0 pp512 PTQ1_0 tg128 PQ2_0 pp512 PQ2_0 tg128
stock prism – 4.0* DeviceLost DeviceLost
pr188 95.42 4.01 141.37 22.74
pr188 + #187 101.19 8.19 142.68 22.72
pr188 + LUT trit decode only (local branch, same idea as #187's decode half, no FWHT fusion) 95.37 4.76 – –

* stock prism has no PQ2_0 Vulkan support at all (device lost on load); its
PTQ1_0 decode already runs the generic dequant path, hence the same ~4.0 t/s.

Why PTQ1_0 stays slow here: the new mul_mat_vec_ptq1_0_q8_1_f32 pipeline is
never created because gfx1013 lacks VK_KHR_shader_integer_dot_product
(int dot: 0), so decode falls back to the generic mul_mat_vec_ptq1_0_f32_f32
shader. To split #187's contribution I tried a local branch with only a table-driven
3^n mod 256 trit decode + vectorized dequantize4 (bit-identical: ppl 2.0268
both, 145/145 backend tests) — that alone gives 4.76 t/s, so most of the 8.19
comes from #187's MUL+FWHT fusion. For int-dot-less devices #187 is the more
important half; I'm not opening a separate PR since #187 already covers the
decode part.

FWIW this board also served PQ2_0 fine from llama-server at ctx 65536 with 2
slots (GTT 8.69 GB), Korean chat and JSON-schema outputs OK.

Thanks for the patch — PQ2_0 on Vulkan is a big win for this board.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

The hand-tuned shader introduces substantial hardware-specific behavior and has only been validated on one GPU and driver combination.

Review effort: Balanced
Findings: None

@renovys

renovys commented Sep 21, 2026

Copy link
Copy Markdown

Follow-up from the BC-250 (gfx1013, RDNA1, int dot: 0, RADV) with more results on your PQ2_0 file.

Two small BC-250-specific improvements, both filed as PRs stacked on this one:

  • vulkan: tune the PQ2_0 mat-vec for AMD BC-250 (gfx1013, RDNA1) #229 — PQ2_0 mat-vec tuned for gfx1013 (4 rows/workgroup, 32-wide, subgroup 32), gated on deviceID 0x13fe && RDNA1 && !coopmat so nothing else changes. llama-bench tg128 22.43 -> 23.93 t/s (+6.7%), pp512 unchanged, PPL identical to the last digit, test-backend-ops 2310/2310.
  • qwen35: apply the Hadamard inverse to the MTP token-embedding lookup #230 — --spec-type draft-mtp (MTP) does not start on the Hadamard-latent PQ2_0-MTP model: graph_mtp fed the raw latent embedding rows into the graph and llama_verify_hadamard_graph throws on the first draft-graph build. Applying the same inverse build_inp_embd() uses fixes it. With the fix, draft-mtp runs at 31.1 t/s (61.2% accept), output byte-identical to spec-off.

Combined, the PQ2_0-MTP model now serves live on this box at ~30.2 t/s single-stream (ctx 65,536, 2 slots), zero DeviceLost/OOM.

Separately, PTQ1_0 stays on the generic path on this int dot: 0 device (decode ~4.8 t/s alone, ~8.2 with #187) and isn't viable here — PQ2_0 is the usable format. Filing that as a tracked issue.

Happy to test further builds on gfx1013.

@bri-prism

Copy link
Copy Markdown
Collaborator

Data point: Intel Arc B390 (Panther Lake Xe3 iGPU), Vulkan / Windows — correctness clean, but the PTQ1_0 decode win is gated off by default

Device line:

Intel(R) Arc(TM) B390 GPU | uma: 1 | fp16: 1 | bf16: 0 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: KHR_coopmat

Intel proprietary Windows driver. This is int dot: 1 and coopmat, which I don't think has been reported on this PR yet — drien's RX 6900 XT has int dot without coopmat, renovys's BC-250 has neither, and rofl0r's Xe iGPU is an earlier part.

Builds: prism @ 422590f5d vs this PR @ 80f71b004, MinGW/ucrt64 g++, Ninja, Release, same flags for both.

Correctness — test-backend-ops test -b Vulkan0, full and unfiltered

prism 422590f5d #188 80f71b004
OK 17187 17434
failing cases 3 3 (same cases)
PTQ1_0 143 OK / 180 not-supported 199 OK / 180 not-supported
PQ2_0 0 OK / 323 not-supported 191 OK / 180 not-supported

The 3 failures are GATED_DELTA_NET(... head_count=4, head_size=128, raw_gates=1), identical on stock prism, so they are pre-existing and unrelated to this PR. No new failures, and PQ2_0 goes from "no Vulkan support at all" to executing.

The finding: the new PTQ1_0 mat-vec never runs on Intel Windows by default

Bonsai 2 27B PTQ1_0, -p 512 -n 128 -ngl 99 -fa 1 -r 3, two ABAB-interleaved rounds:

build tg128 (r1 / r2) pp512 (r1 / r2)
prism 1.59 / 1.57 206.6 / 169.9
#188, default 1.55 / 1.57 154.4 / 169.3

No change. The cause is in ggml_vk_should_use_mmvq(): the INTEL_XE2 early-out whitelists only Q2_0 / Q2_K / Q3_K / Q6_K, and immediately below it

if (device->driver_id == vk::DriverId::eIntelProprietaryWindows) {
    // Intel Windows proprietary driver MMVQ performance for !Q2/Q3/Q6 is worse than fp16,
    return false;
}

returns before the per-type switch. So on this driver PTQ1_0 (and PQ2_0) never reach mul_mat_vecq_ptq1_0.comp.

Forcing the path, same models and flags:

build tg128 pp512
prism + GGML_VK_FORCE_MMVQ=1 1.62 ± 0.05 217.1
#188 + GGML_VK_FORCE_MMVQ=1 13.16 ± 0.01 215.6

8.1x decode, in line with the 8x in the PR description. The prism row with the same env var is the control — it stays at 1.62, so the win is this PR's kernel and not the flag.

Suggestion: let PTQ1_0 / PQ2_0 past that carve-out on Intel (e.g. add them to the INTEL_XE2 whitelist above it). The comment and the two issues it links predate both types, and on this part the blanket opt-out is costing exactly the 8x this PR is for.

PQ2_0 needs no flag

PQ2_0 goes through the dequant mat-vec this PR adds rather than mmvq, so it is unaffected by the gate: pp512 314.3, tg128 11.56 (with GGML_VK_FORCE_MMVQ=1: 294.4 / 11.63 — no meaningful difference). On stock prism this model has no Vulkan path at all.

Caveats

Single machine, and an iGPU on UMA, so "VRAM" is system DRAM at the same bandwidth. pp512 on this box carries roughly 20% round-to-round spread from page-cache state (visible in the two prism rows), so I would not read the prefill column as a regression either way; tg128 spread is under 1% and is where the signal is.

@bri-prism

Copy link
Copy Markdown
Collaborator

Followed up on the mmvq gate above with a patch rather than just a report: #238, stacked on this branch.

It adds PTQ1_0 to the INTEL_XE2 whitelist in ggml_vk_should_use_mmvq() so it gets past the eIntelProprietaryWindows opt-out — seven lines, five of them comment. On the default path with no env overrides, this branch plus that commit gives tg128 1.56 -> 13.02 t/s on Arc B390 with Ternary-Bonsai-2-27B-PTQ1_0, i.e. your 8x lands on Intel Windows too instead of being unreachable there.

PQ2_0 is deliberately left out of that whitelist — measured 11.56 vs 11.63 with and without mmvq, so it loses nothing by staying on the dequant mat-vec this PR adds.

Full test-backend-ops is unchanged against this branch (17434 OK, same 3 pre-existing GATED_DELTA_NET failures, PTQ1_0 199 OK / 0 FAIL). One side effect worth knowing: because the gate was blocking mmvq here, this branch's own test run never exercised mul_mat_vecq_ptq1_0.comp on this device — the 199 PTQ1_0 cases all went through the dequant path. With the whitelist change they route through your new shader and still pass, so #238 doubles as first Intel coverage of that kernel.

Entirely your call whether to fold it into this PR or keep it separate — happy either way, and happy to rebase it if this one moves.

@bri-prism

Copy link
Copy Markdown
Collaborator

Thanks for this work, @alhnesn. It has landed through #238, which was stacked on this branch and merged on 09-21 as 3ae4f51: your integer-dot PTQ1_0 mat-vec shader and the PQ2_0 support are on prism byte for byte, and every line this PR adds to ggml-vulkan.cpp is there too. It also shipped in the new release prism-b10735-842b188. Since merging this now would add nothing, I'm closing it as superseded by #238. Thanks again, this is what made PTQ1_0 decode fast on Vulkan.

@bri-prism bri-prism closed this Sep 24, 2026
@alhnesn

alhnesn commented Sep 24, 2026

Copy link
Copy Markdown
Author

Thanks @bri-prism, sorry I was away for a few days. Glad it landed through #238, and thanks for testing it on Intel and getting it working there too

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants