Skip to content

ggml-cpu : move repacked weight rows to the NUMA node of the thread that reads them - #251

Open
lenny76 wants to merge 1 commit into
PrismML-Eng:prismfrom
lenny76:cpu-repack-numa-local-pages
Open

lenny76 wants to merge 1 commit into
PrismML-Eng:prismfrom
lenny76:cpu-repack-numa-local-pages

Conversation

@lenny76

@lenny76 lenny76 commented Sep 23, 2026

Copy link
Copy Markdown

Overview

With a NUMA strategy (--numa distribute or isolate) the repack forward_mul_mat disables dynamic chunking, so every thread always reads the same row slice of a weight, and threads are bound to nodes. The repacked weights, however, are written by the loading thread, so all of them end up on one node: half of the threads read remote memory and the whole matmul waits for them.

This PR makes each thread move the pages of its row slice of a repacked weight to its own node (move_pages) the first time it computes that weight. It runs once per weight and thread count, only on Linux, only when ggml_is_numa(). Nothing changes without --numa.

Additional information

2x Xeon Gold 6262 (Cascade Lake, 24 cores per socket), Linux, GCC 14, -DGGML_NATIVE=ON. Page placement of the repacked weights during decode (/proc/PID/numa_maps), PQ2_0: before 6.34 GiB on one node, after 3.21 / 3.14 GiB.

Ternary-Bonsai-2-27B-PQ2_0.gguf, base prism + #245 + #246 + #249 + #250, llama-bench -ngl 0 --numa distribute -p 0 -n 64 -r 3, two rounds:

memory policy threads base tg64 PR tg64
default 24 5.37 / 5.33 6.75 / 6.69
default 48 5.34 / 5.32 6.18 / 6.18
numactl --interleave=all 24 6.52 / 6.47 6.58 / 6.64
numactl --interleave=all 48 6.51 / 6.45 6.64 / 6.61
no --numa 24 4.79 4.81

llama-completion --numa distribute -t 24: eval 4.09 -> 6.97 t/s (two runs: 6.97, 7.01). Load time grows by about 1 s (the pages move during the warmup decode); prompt eval time is unchanged.

The same patch on upstream master 633733d0a with Qwen3-8B-Q4_K_M (Q4_K is repacked on AVX2), --numa distribute -p 512 -n 64:

memory policy threads master tg64 PR tg64
default 8 6.07 / 6.05 8.44
default 24 6.49 / 6.51 9.38
default 48 6.04 / 6.16 8.59
numactl --interleave=all 8 7.27 / 7.26 8.35
numactl --interleave=all 24 9.37 / 9.38 9.38

pp512 is unchanged (within noise) in all cases. Greedy completions are identical with and without the PR.

Limits: if prompt processing and generation use different thread counts, the pages are moved again for the second one. With --numa numactl threads are not bound per node by ggml, so the placement follows wherever each thread runs at its first use.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES. Claude Code (Anthropic) wrote the patch and ran the measurements above on my machine, under my direction.

🤖 Generated with Claude Code

…eads them

With a NUMA strategy (--numa distribute / isolate) the repack mul_mat
uses a static row split and threads are bound to nodes, but the
repacked weights were written by the loading thread, so all of them
sit on one node. Each thread now moves the pages of its row slice of
a weight to its own node the first time it computes that weight
(move_pages, once per weight and thread count, Linux only).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bri-prism

Copy link
Copy Markdown
Collaborator

Built on Windows (MSYS2 UCRT64 GCC 16.2, GGML_NATIVE=ON) merged onto prism @ 3b19c377d: compiles cleanly, so the Linux-only guard works on MinGW. I couldn't test the NUMA behaviour itself: this is a single-socket Intel laptop running Windows. No functional test beyond the build.

Tested with Claude Code.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants