Skip to content

qwen35: allow Vulkan for rows-mode GDN state + Vulkan 890M recipe doc - #299

Open
MrFadiAi wants to merge 1 commit into
PrismML-Eng:prismfrom
MrFadiAi:qwen35-vulkan-rows-recipe
Open

MrFadiAi wants to merge 1 commit into
PrismML-Eng:prismfrom
MrFadiAi:qwen35-vulkan-rows-recipe

Conversation

@MrFadiAi

Copy link
Copy Markdown

Split-out part 3 of #187, rebased on current prism, per review feedback.

Scope: one allowlist line + docs.

  • qwen35.cpp: gdn_state_rows_dev_ok no longer excludes Vulkan. The gated_delta_net_rows pipelines (see the bf16-SSM-state PR) run natively on Vulkan; excluding it forced a per-layer CPU round-trip (~192 extra kernel launches/token) on hybrid models.
  • docs/vulkan-890m-bonsai2-recipe.md: complete Vulkan 890M speed recipe for Bonsai-2-27B — build config, launch flags, env vars, the measured optimization ladder (1.85 → 14 t/s), and every pitfall we hit on the way.

Measured, Radeon 890M (RDNA 3.5), Bonsai-2-27B Q2_0-fork + MTP n-max 1 + LLAMA_SSM_BF16_STATE=1:

  • single-stream: 12.7 → 13.6–14.1 t/s, 5/5 generation gates (spec-acceptance 0.84)

… doc

- qwen35.cpp: gdn_state_rows_dev_ok no longer excludes Vulkan. The
  gated_delta_net_rows pipelines run natively on Vulkan; excluding it
  forced a per-layer CPU round-trip (~192 extra launches/token).
- docs: complete Vulkan 890M speed recipe for Bonsai-2-27B.

Measured, Radeon 890M, Bonsai-2-27B Q2_0-fork + MTP n-max 1
+ LLAMA_SSM_BF16_STATE=1: single-stream 12.7 -> 13.6-14.1 t/s, 5/5 gates.

Co-authored-by: Hermes Agent <noreply@nousresearch.com>
@MrFadiAi
MrFadiAi force-pushed the qwen35-vulkan-rows-recipe branch from a3143a7 to 5018ce6 Compare September 30, 2026 12:59

@bri-prism bri-prism left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for splitting this out. A few things before this can go in:

  1. Merge order. On prism today, Vulkan's supports_op for GATED_DELTA_NET still returns false when src[6] is set. If this lands before #298, every recurrent layer's GDN op falls back to the CPU on Vulkan when speculation is on. Measured on an Arc B390 (27B PQ2_0 with a 2B draft, n_max 1): speculative decode drops from about 13 t/s to 1.1 t/s, and one prose prompt produced degenerate output (the model repeats the prompt). Plain decode is unaffected. Could you either fold this change into #298 or mark it as depending on #298?

  2. Default F32 state. All the numbers here use LLAMA_SSM_BF16_STATE=1. With speculative decoding on (which is what enables rows mode), this change turns rows mode on for Vulkan, and most users will be on the default F32 state. Please share a run with the env var unset, plus test-backend-ops -o GATED_DELTA_NET output with rows mode on Vulkan.

  3. Stale comment. The comment above gdn_state_rows_env in qwen35.cpp still says rows mode runs on CPU and Metal only. Please update it along with the allowlist.

  4. Coverage. Every result so far is from the 890M. Do you have a run on any other Vulkan device, such as a discrete AMD card or Intel?

On the doc, I'd split it into its own PR. A few requests for that version:

  • Please remove section 3 (how the model files are built and where the tensors come from). That isn't something we want documented in this repo.
  • The references to building from the #187 branch and the #187 commit hashes are out of date since the split.
  • The env var is spelled two ways, GGML_SSM_BF16_STATE and LLAMA_SSM_BF16_STATE.
  • The leaner-quant results disagree: the table says IQ3, finding 3 says PTQ1_0.
  • The 1.85 t/s baseline uses a different build and model file, so the 7.5x figure doesn't isolate these changes. A before/after on the same file would be clearer.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation model

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants