Conversation
…eads them With a NUMA strategy (--numa distribute / isolate) the repack mul_mat uses a static row split and threads are bound to nodes, but the repacked weights were written by the loading thread, so all of them sit on one node. Each thread now moves the pages of its row slice of a weight to its own node the first time it computes that weight (move_pages, once per weight and thread count, Linux only). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Collaborator
|
Built on Windows (MSYS2 UCRT64 GCC 16.2, Tested with Claude Code. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
With a NUMA strategy (
--numa distributeorisolate) the repackforward_mul_matdisables dynamic chunking, so every thread always reads the same row slice of a weight, and threads are bound to nodes. The repacked weights, however, are written by the loading thread, so all of them end up on one node: half of the threads read remote memory and the whole matmul waits for them.This PR makes each thread move the pages of its row slice of a repacked weight to its own node (
move_pages) the first time it computes that weight. It runs once per weight and thread count, only on Linux, only whenggml_is_numa(). Nothing changes without--numa.Additional information
2x Xeon Gold 6262 (Cascade Lake, 24 cores per socket), Linux, GCC 14,
-DGGML_NATIVE=ON. Page placement of the repacked weights during decode (/proc/PID/numa_maps), PQ2_0: before 6.34 GiB on one node, after 3.21 / 3.14 GiB.Ternary-Bonsai-2-27B-PQ2_0.gguf, baseprism+ #245 + #246 + #249 + #250,llama-bench -ngl 0 --numa distribute -p 0 -n 64 -r 3, two rounds:numactl --interleave=allnumactl --interleave=all--numallama-completion --numa distribute -t 24: eval 4.09 -> 6.97 t/s (two runs: 6.97, 7.01). Load time grows by about 1 s (the pages move during the warmup decode); prompt eval time is unchanged.The same patch on upstream master
633733d0awithQwen3-8B-Q4_K_M(Q4_K is repacked on AVX2),--numa distribute -p 512 -n 64:numactl --interleave=allnumactl --interleave=allpp512 is unchanged (within noise) in all cases. Greedy completions are identical with and without the PR.
Limits: if prompt processing and generation use different thread counts, the pages are moved again for the second one. With
--numa numactlthreads are not bound per node by ggml, so the placement follows wherever each thread runs at its first use.Requirements
🤖 Generated with Claude Code