Skip to content

metal: opt-in Metal 4.1 direct Q2 operand experiment - #297

Draft
bri-prism wants to merge 2 commits into
prismfrom
metal41-direct-optin
Draft

bri-prism wants to merge 2 commits into
prismfrom
metal41-direct-optin

Conversation

@bri-prism

@bri-prism bri-prism commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Overview

Adds a disabled-by-default direct cooperative-operand Q2_0/PQ2_0 matrix path for comparison on M5. Set GGML_METAL_Q2_DIRECT_MAX=16; the minimum width defaults to 8. Enabling it requires macOS 27, M5, TensorOps, embedded Metal sources, and a successful Metal 4.1 compiler probe. The shader is excluded and existing dispatch retained when the gate is off.

Draft: this is slower than the current branch's PQ2_0 few-row path from #284/#286 and is not recommended for default selection. The earlier staged-kernel baseline has been superseded. The candidate preserves the activation operand type and also covers Q2_0 and generic matrix batch strides; these differences need further evaluation before deciding whether to retain the implementation.

Validation

  • Release build and focused Q2_0/PQ2_0 MUL_MAT backend checks passed with opt-in off/on. Unsupported reference combinations are skipped.
  • M5 Pro, macOS 27.0.1, Xcode 27: three alternating pairs, three repetitions each, depth 128, FA off. Current/candidate median block times: width 8 = 59.18/95.79 ms; width 16 = 72.57/114.53 ms.
  • git diff --check passed. Changed C++/Metal ranges formatted; Objective-C retained local style because the repository formatter excludes that language.
  • Full CI, old-device runtime fallback, current-branch model-logit parity, broader accuracy and speculative end-to-end validation remain outstanding. No defaults changed.

Usage and limitations: docs/development/metal-direct-q2.md.

Follow-up findings

Four-way K splitting passed the backend oracle and reduced the candidate cost, but still measured 1.331x/1.345x current-default latency at widths 8/16. Two independent variants of that split also lost: half activation conversion (1.531x/1.555x) and original-type packed-word sharing (1.511x/1.462x). Each used three alternating pairs with three repetitions. Rejected follow-up variants were removed; no extra knobs added.

The unchanged existing kernel under Metal 4.0 versus 4.1 differed by less than 0.2%, ruling out a meaningful compiler-language benefit in this screen. This remains a draft comparison path, not a proposed default replacement.

Requirements

  • Contributing guidelines reviewed; submitted as a draft for maintainer review.
  • AI usage disclosure: YES - implementation, local experiments, and draft preparation were assisted by Codex. The submitting contributor remains responsible for review and maintenance.

@github-actions github-actions Bot added documentation Improvements or additions to documentation ggml Apple Metal labels Sep 30, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Apple Metal documentation Improvements or additions to documentation ggml

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant