Conversation
… doc - qwen35.cpp: gdn_state_rows_dev_ok no longer excludes Vulkan. The gated_delta_net_rows pipelines run natively on Vulkan; excluding it forced a per-layer CPU round-trip (~192 extra launches/token). - docs: complete Vulkan 890M speed recipe for Bonsai-2-27B. Measured, Radeon 890M, Bonsai-2-27B Q2_0-fork + MTP n-max 1 + LLAMA_SSM_BF16_STATE=1: single-stream 12.7 -> 13.6-14.1 t/s, 5/5 gates. Co-authored-by: Hermes Agent <noreply@nousresearch.com>
a3143a7 to
5018ce6
Compare
bri-prism
left a comment
There was a problem hiding this comment.
Thanks for splitting this out. A few things before this can go in:
-
Merge order. On
prismtoday, Vulkan'ssupports_opfor GATED_DELTA_NET still returns false whensrc[6]is set. If this lands before #298, every recurrent layer's GDN op falls back to the CPU on Vulkan when speculation is on. Measured on an Arc B390 (27B PQ2_0 with a 2B draft, n_max 1): speculative decode drops from about 13 t/s to 1.1 t/s, and one prose prompt produced degenerate output (the model repeats the prompt). Plain decode is unaffected. Could you either fold this change into #298 or mark it as depending on #298? -
Default F32 state. All the numbers here use
LLAMA_SSM_BF16_STATE=1. With speculative decoding on (which is what enables rows mode), this change turns rows mode on for Vulkan, and most users will be on the default F32 state. Please share a run with the env var unset, plustest-backend-ops -o GATED_DELTA_NEToutput with rows mode on Vulkan. -
Stale comment. The comment above
gdn_state_rows_envinqwen35.cppstill says rows mode runs on CPU and Metal only. Please update it along with the allowlist. -
Coverage. Every result so far is from the 890M. Do you have a run on any other Vulkan device, such as a discrete AMD card or Intel?
On the doc, I'd split it into its own PR. A few requests for that version:
- Please remove section 3 (how the model files are built and where the tensors come from). That isn't something we want documented in this repo.
- The references to building from the #187 branch and the #187 commit hashes are out of date since the split.
- The env var is spelled two ways,
GGML_SSM_BF16_STATEandLLAMA_SSM_BF16_STATE. - The leaner-quant results disagree: the table says IQ3, finding 3 says PTQ1_0.
- The 1.85 t/s baseline uses a different build and model file, so the 7.5x figure doesn't isolate these changes. A before/after on the same file would be clearer.
Split-out part 3 of #187, rebased on current prism, per review feedback.
Scope: one allowlist line + docs.
qwen35.cpp:gdn_state_rows_dev_okno longer excludes Vulkan. Thegated_delta_net_rowspipelines (see the bf16-SSM-state PR) run natively on Vulkan; excluding it forced a per-layer CPU round-trip (~192 extra kernel launches/token) on hybrid models.docs/vulkan-890m-bonsai2-recipe.md: complete Vulkan 890M speed recipe for Bonsai-2-27B — build config, launch flags, env vars, the measured optimization ladder (1.85 → 14 t/s), and every pitfall we hit on the way.Measured, Radeon 890M (RDNA 3.5), Bonsai-2-27B Q2_0-fork + MTP n-max 1 +
LLAMA_SSM_BF16_STATE=1: