Skip to content

feat: rank quantisations by an explicit ladder - #5

Merged
olegshulyakov merged 1 commit into
mainfrom
feat/quant-ladder
Aug 31, 2026
Merged

olegshulyakov merged 1 commit into
mainfrom
feat/quant-ladder

Conversation

@olegshulyakov

Copy link
Copy Markdown
Collaborator

The quality order now lives in QuantLadder in plan.go, and measure, validate and pick all read the band from there instead of each keeping their own copy.

  • Add UD-Q6_K_XL. It ships in all 71 repos, but it was never in the band and the filename pattern could not match it either, so the best in-band quant was silently skipped on every model. It now replaces plain Q6_K wherever it fits the budget; Q6_K stays for the tiers where it does not.
  • Drop UD-Q6_K, UD-Q5_K_M and UD-Q4_K_M. Every repo that still carries one of those old names also ships the larger _XL sibling, so under an explicit ladder they can never be picked.
  • Rank by ladder position instead of measured size. Size flips the dynamic and plain families between models - Qwen3-30B-A3B ships UD-Q4_K_XL at 16.48 GiB against Q4_K_M at 17.28, gemma-3-27b-it the same pair at 15.67 against 15.41 - so on half the catalogue it ranked a plain quant above the dynamic quant of the same bit level.
  • Fix quantOf reading "UD-Q4_K_M.gguf" as plain Q4_K_M. An alternation of the band alone matches the in-band tag inside the longer name, and the trailing ".gguf" anchor does not help because that tag really does end the name. So narrowing the band turned those files into fake measurements of a quant their repo does not publish. The tag is now captured generically and compared against the ladder, with a test for each collision.
  • Give each GGUF header read its own deadline. A connection to the HF CDN that stopped responding mid-handshake never returned an error, so the retry never fired and measure parked all eight workers at 0% CPU, which also blocked every finished measurement behind it in the save loop.

Tier membership is unchanged at 22/36/62/71/71 models. There are seven more sections in total, because quality and balanced now land on different quants more often and so merge less.

What does this PR do?

Why?

How to test

Checklist

  • Changes are focused and minimal
  • Code is readable and commented where needed
  • Tests pass (if applicable)
  • README or docs updated if behavior changed
  • No unrelated changes included

Related issues

The quality order now lives in QuantLadder in plan.go, and measure, validate and
pick all read the band from there instead of each keeping their own copy.

- Add UD-Q6_K_XL. It ships in all 71 repos, but it was never in the band and the
  filename pattern could not match it either, so the best in-band quant was
  silently skipped on every model. It now replaces plain Q6_K wherever it fits
  the budget; Q6_K stays for the tiers where it does not.
- Drop UD-Q6_K, UD-Q5_K_M and UD-Q4_K_M. Every repo that still carries one of
  those old names also ships the larger _XL sibling, so under an explicit ladder
  they can never be picked.
- Rank by ladder position instead of measured size. Size flips the dynamic and
  plain families between models - Qwen3-30B-A3B ships UD-Q4_K_XL at 16.48 GiB
  against Q4_K_M at 17.28, gemma-3-27b-it the same pair at 15.67 against 15.41 -
  so on half the catalogue it ranked a plain quant above the dynamic quant of
  the same bit level.
- Fix quantOf reading "UD-Q4_K_M.gguf" as plain Q4_K_M. An alternation of the
  band alone matches the in-band tag inside the longer name, and the trailing
  ".gguf" anchor does not help because that tag really does end the name. So
  narrowing the band turned those files into fake measurements of a quant their
  repo does not publish. The tag is now captured generically and compared
  against the ladder, with a test for each collision.
- Give each GGUF header read its own deadline. A connection to the HF CDN that
  stopped responding mid-handshake never returned an error, so the retry never
  fired and measure parked all eight workers at 0% CPU, which also blocked every
  finished measurement behind it in the save loop.

Tier membership is unchanged at 22/36/62/71/71 models. There are seven more
sections in total, because quality and balanced now land on different quants
more often and so merge less.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@olegshulyakov
olegshulyakov merged commit 0423a5f into main Aug 31, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant