Repository navigation
Free a level's staging once its upload releases it, and report app memory - #1079
Merged
Merged
Conversation
The Metal HUD's app memory creeps up while GTA IV plays, and the layer's own figures could not say where it goes. With the address-space watch at debug, page boxes grew through the unattributed "other" holder, from 0 to 544 MiB in two minutes and then about 26 MiB a minute to 631 MiB, while textures (about 5000), mip data (about 710 MiB), staging (about 299 MiB), vertex/index backing and the parked pool (128 MiB cap) stayed flat. Free 32-bit address space fell about 80 MiB a minute late in the run, three times the page-box growth, and nothing in the perf summary covered the rest. The perf summary now reads three gauges once per window, only when it is about to emit, so a build without PERF=1 pays nothing and a PERF=1 frame pays nothing outside the emitting one: - process_footprint_bytes, the process's phys_footprint through proc_pid_rusage, the figure the HUD shows as app memory; - metal_allocated_bytes, the MTLDevice's currentAllocatedSize; - tex_staging_wrapped_bytes, the staging under the encoder's cached per-level bytesNoCopy wrappers, each of which keeps its upload lease and so its guest pages alive. They ride CacheSizes beside the cache counts, get footprint, metal alloc and wrapped rows in the grid, and are left out of a window closed before their sample arrived, as the fault keys are. The bench classifies the two process-wide gauges as info, since they count Wine, the driver and what earlier benchmarks of a round left behind; the wrapper gauge keeps the bytes rule. The address-space watch names the texture staging only upload leases keep after the texture let go of it, which it counted as other. The sample walks the device's retained page leases and counts an allocation only when every reference to it is a lease's, so staging a texture still holds stays under texture staging. Beside the page boxes it logs what d3d9.dll's snmalloc has committed, from the allocator's own statistics, two atomic loads.
The encoder caches one bytesNoCopy MTLBuffer per staging level and keeps the native owner of the level's pages in the slot as the wrapper's keepalive. That owner was adopted from the upload lease that first delivered the pages, and the lease completes only when the last native owner drops, so the PE side keeps the pages for as long as the wrapper stays cached. The slot was emptied only when the level's backing changed or the texture was destroyed. A DEFAULT-pool static level releases its staging once its upload is emitted: the upload record carries release_staging, the encoder's answer reaches the API thread at the next Present, and release_emitted_staging drops the texture's copy. The cached wrapper kept the pages alive anyway, charged to no holder. In a GTA IV run (v0.11.0-91-ge99fa026, mem_watch at debug) that "other" share grew from 0 to 544 MiB in two minutes and then about 26 MiB a minute to 631 MiB, while texture staging held at about 299 MiB and the texture count at about 5000. An emitted upload that carries release_staging now takes the level's wrapper out of the cache and parks it on the resource retention queue behind the current submission, the way a backing change already did, keepalive included. Its destroy waits for both the upload and draw counters, the keepalive drops after the destroy, and the lease then completes, so the pages are freed or parked in the staging lane once the GPU retires that upload. The textures destroys row now counts these retirements too. The PE side applies the answer at the next Present and only when no newer upload of the level overtook it, so a DEFAULT-pool level rewritten every frame keeps its staging, and retiring on every answer would create and destroy a wrapper per level per frame over the same backing. The slot an answer empties therefore keeps the backing's address and length; when the next upload wraps that same backing again, the new wrapper is marked kept and later answers leave it cached. That bounds the churn to one extra wrapper per backing. A kept wrapper still retires at a backing change or at texture destroy, so a level that is rewritten for a while and then released keeps the old behaviour of holding its pages until then. Carrying the PE side's actual drop back to the encoder would retire exactly, but needs a new frame operation for it. The perf summary counts the churn directly: tex_wrapper_create_total and tex_wrapper_retire_total, the wrappers created and queued for destroy per window, with a churn row under the textures wrapped row. The bench reads them as noisy per-frame rates.
athei
added a commit
that referenced
this pull request
Oct 7, 2026
## Symptoms The v0.12.0 release bench (`make bench-ab BASE=v0.11.0 BENCH_SET=full`) hit three problems in the harness itself, separate from what the layer changed. First, it stopped with exit 2 and no verdict: `metric perf.metal_allocated_bytes is in some rounds of the cand leg and not in others`. The candidate rounds lacked the memory gauges in query_poll_wow round 1, streaming round 2 and shader_stutter_offscreen3 round 0. #1079 added `process_footprint_bytes`, `metal_allocated_bytes` and `tex_staging_wrapped_bytes` to the `perf-kv` line and documented that a window closed before their sample arrived leaves them out, but the comparer only allowed the fault counts and the pool's GPU copies to come and go. Second, once the comparer gave a verdict, `shader_stutter_k1`'s `frame.spikes` read as a regression, 1 to 16. A frame counted as a spike when it took over twice the median, with no absolute floor. With pipeline compiles off the API thread the candidate's median is about 31 us, so a 62 us frame counted, while its slowest frame took 0.47 ms against 46 ms on v0.11.0. Third, `cold_start`'s `prewarm.shaders`, an exact count, read 98 on v0.11.0 and 92 on the candidate and failed the run. #1050 stopped keying `D3DRS_SPECULARENABLE` into the vertex shader of an unlit fixed-function draw, and the benchmark toggles specular on 12 unlit combinations, so six pairs now share one shader. Nothing goes unprepared: both legs pre-warm 196 pipelines with none failed or skipped, and the candidate pre-warms all 92 shaders its cache holds. ## Changes `OPTIONAL_METRICS` in `unix/e2e/src/bench/compare.rs` now lists `perf.process_footprint_bytes`, `perf.metal_allocated_bytes` and `perf.tex_staging_wrapped_bytes`. A round without them is reported as incomplete with a note and never judged, the same as the fault counts. The module and constant docs name the six keys, matching the list in `docs/ARCHITECTURE.md`. The wrapper churn counts (`tex_wrapper_create_total`, `tex_wrapper_retire_total`) are written in every window, so they stay required. This carries the limitation the fault counts already have. Any optional key missing from a whole leg, or from one round of a leg, becomes an incomplete row, and an incomplete row never fails a run. A candidate that stopped writing the gauges in every round while the base has them would read as incomplete, not REMOVED. `tex_staging_wrapped_bytes` is a judged bytes row, and it goes unjudged in any benchmark where one round misses the sample, which was 3 of the 13 benchmarks in the release run. A unit test pins this behaviour so a change to it is deliberate. `windows/tests/tests/e2e/bench.rs` gets `FRAME_SPIKE_FLOOR` at 1 ms beside the 50 us `SPIKE_FLOOR` of the call timers, and `FrameClock::spikes` replaces `FrameClock::over`: a frame is a spike only when it is over both twice the median and the floor, the rule `CallTimes::spikes` uses for single calls. The benchmarks present without waiting for the display, so twice a median of tens of microseconds is a preempted thread or a timer. A pipeline compile on the API thread takes milliseconds (one new shader a frame put v0.11.0's median at 11 ms), so 1 ms still counts every frame that waited for one, and it is 6 % of a 60 Hz frame. With a median above 0.5 ms the limit is twice the median as before. `frame.spikes` keeps its name: both legs of `make bench-ab` run the candidate's benchmark binary, so both count by the same rule, and the metric's unit, direction and class do not change. The shader-stutter report line and the `COVERAGE.md` rows name the floor. The harness accepts an exact change per run only (`ACCEPT`, `--accept`) and keeps no list of expected changes, so the `prewarm.shaders` change is documented rather than recorded in code: the benchmark section of `CONTRIBUTING.md` now says that a comparison against v0.11.0 or older takes `ACCEPT=prewarm.shaders`, why, and that the acceptance belongs to that run alone, so a later change to the count still fails. ## Alternatives Waiting in the encoder until the memory sample arrives would keep the gauges in every window, but it changes the layer to suit the harness, and the line's contract already says a consumer treats a missing key as not measured. A list of expected exact changes keyed by base version would let a run accept `prewarm.shaders` without naming it, but nothing else would use it, and a per-run `ACCEPT` already keeps a later change to the count visible. ## Verification `memory_gauges_a_candidate_window_left_out_are_reported_not_judged` in `unix/e2e/src/bench/compare/tests.rs` writes a base leg without the gauges and a candidate leg whose round 1 lacks them, and checks that a verdict comes back with the three rows incomplete and the churn rows judged. It fails on main with the error the release bench printed. `memory_gauges_the_candidate_drops_from_every_round_are_incomplete_not_removed` gives the base the gauges in every round and the candidate none, and checks that the three rows are incomplete, none is removed and the run passes. `frame_spikes_below_the_floor_are_not_counted` and `frame_spikes_over_the_floor_and_twice_the_median_are_counted` in `windows/tests/tests/e2e/bench_clock.rs` test the rule on fixed frame times: the release run's candidate frames count no spike, and a frame over the floor, or over twice a median of milliseconds, counts. They are in the end-to-end binary, which runs under Wine; for this PR they were built and linted by `make check`, not run. `make bench-compare` on a copy of the failed release run now gives a verdict instead of exit 2: FAIL, 13 benchmarks, 5 regressed, 93 improved, 1 exact changed, 60 incomplete, 1 of 2 shapes changed. With `ACCEPT=prewarm.shaders` the exact change reads "changed, accepted". The stored run's `frame.spikes` was counted by the old benchmark binary, so it still shows the old count; the floor takes effect on the next `make bench-ab`. The other regressions and the shape change are the release comparison's own findings and are not touched here. `make fmt`, `make check` and `make test-unit` pass. ## Rules check No mutable static, config key, wire field, dependency, derive or lint suppression is added. `FRAME_SPIKE_FLOOR` is a `const`, next to the `SPIKE_FLOOR` it mirrors; the frame rule repeats the shape of `CallTimes::spikes` over a different input (frame times instead of sampled calls) and a different floor, and `FrameClock::over` goes away rather than staying beside it. The comparer change extends an existing constant. The comparer's tests live in `compare/tests.rs` as `docs/CONVENTIONS.md` requires, and the frame clock's in `bench_clock.rs`, whose `COVERAGE.md` row is extended; `make check` (audit) enforces both. Doc comments keep the title, blank line, body shape and the 100-column limit, checked by `make check` (audit and doc).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Symptoms
In GTA IV the Metal HUD's app memory, the process footprint, creeps up during play. A run of v0.11.0-91-ge99fa026 with
RUST_LOG=mtld3d::d3d9::mem_watch=debug(GTAIV-12703.log) shows the page boxes growing through the unattributed holder:othergoes from 0 to 544 MiB in the first two minutes and then about 26 MiB a minute to 631 MiB, while the live textures (about 5000), their mip data (about 710 MiB), texture staging (about 299 MiB), the vertex/index backing and the parked pool (at its 128 MiB cap) stay flat. Free 32-bit address space falls about 80 MiB a minute late in the run, about three times the page-box growth, and nothing in the perf summary covers the process's footprint, the Metal device's allocations ord3d9.dll's heap, so the rest of the growth had no figure at all.Cause
The encoder caches one
bytesNoCopyMTLBufferper texture staging level (MipStagingBufferinunix/unix/src/encoder.rs) and keeps the native owner of the level's pages in the slot as the wrapper's keepalive. That owner is theArc<PageBox>adopted from the upload lease that first delivered the pages, and a lease completes only when its last native owner drops, so while the wrapper is cached the PE side'sGuestPageLeasekeeps its owner of the pages too. The slot was emptied only when the level's backing changed or the texture was destroyed.A DEFAULT-pool static level releases its staging once its upload has been emitted:
schedule_uploadsetsrelease_stagingon the upload record, the encoder's answer reaches the API thread at the nextPresent, andrelease_emitted_stagingdrops the texture's copy. The cached wrapper kept the pages alive anyway, until the texture was destroyed, and since the texture no longer counted them the address-space watch charged them toother. The mechanism is read from the source and the existing unit test that pins it (cached_native_ownership_keeps_original_owner_aliveinwindows/core/src/guest_pages/tests.rs). A pin that lasts until texture destroy explains a plateau proportional to the released levels of the live textures; it did not obviously explain a steady creep of about 26 MiB a minute while the texture count stays flat. The GTA IV run with this change (Verification) settles the page-box part:otherstays at 0 and the page boxes stay flat. The process footprint still grows, and the new figures place that growth outside the layer's allocations.Changes
The fix: an emitted upload that carries
release_stagingtakes the level's wrapper out of the cache and parks it on the resource retention queue behind the current submission, the same path a backing change already took (nowpark_staging_wrapperovertake_staging_wrapper). Its destroy waits for both the upload and the draw counter to pass that submission, the keepalive drops after the destroy, and the lease then completes, so the pages are freed or parked in the staging lane once the GPU retires that upload.The PE side applies the answer at the next
Present, and only when no newer upload of the level overtook it (is_latest_upload), so a DEFAULT-pool level rewritten every frame (anUpdateTexturethen a draw) carriesrelease_stagingon every upload yet keeps its staging. Retiring on every answer would then create and destroy a wrapper per level per frame over the same backing. The slot an answer empties therefore keeps the backing's address and length, and when the next upload wraps that same backing again (MipStagingBuffer::created), the new wrapper is markedkept_after_releaseand later answers leave it cached (take_released_staging_wrapper). The churn is one extra wrapper per backing. A kept wrapper still retires at a backing change or at texture destroy, so a level rewritten for a while and released afterwards holds its pages until then, as every level did before this change. The two new counterstex_wrapper_create_totalandtex_wrapper_retire_total, with achurnrow under the textureswrappedrow, show the wrappers created and queued for destroy per window, so a bench run or a game log shows the churn directly; the bench reads them as noisy per-frame rates.Nothing crosses the boundary that did not before: the encoder already reads
release_stagingfrom the upload record. The texturesdestroysrow andtex_destroy_totalnow include these wrapper destroys; the key's meaning (texture-lifecycle destroys, staging wrappers included) is unchanged, but its count moves earlier, from texture release to upload retirement.Three new
perf-kvgauges follow existing paths.process_footprint_bytesis the process'sphys_footprint, read withproc_pid_rusage(ri_phys_footprint, the same ledger valueTASK_VM_INFOreports), which is what the HUD shows as app memory.metal_allocated_bytesis theMTLDevice'scurrentAllocatedSize, which until now was read only when a texture was refused.tex_staging_wrapped_bytesis the padded staging under the encoder's cached per-level wrappers.The encoder reads the three once per perf window, only on the frame whose summary is about to emit, and carries them in
CacheSizesasmemory: Option<MemoryGauges>, beside the cache counts the summary already samples there. APERF=1build pays for them once every two seconds, a build without it never calls them. The grid getsfootprintandmetal allocrows under the fault row and awrappedrow in Resources (textures), all throughres_row. A window closed between the encoder's due check and the emit check gets no sample: the grid printsn/aand the line leaves the keys out, as it does with the fault keys.In the address-space watch (
windows/d3d9/src/device/mem_watch.rs) the page boxes gain anupload leasesholder: texture staging that only the device's upload leases still keep. The sample walks the device's retained page leases (PacketRetirement::tally_page_leases) and counts an allocation only when every reference to it is a lease's (mtld3d_core::guest_pages::LeaseOnlyPages), so staging a texture still holds stays under texture staging andotheris left with read-back pages of a call in progress and anything not yet named. Beside the page boxes the line logsd3d9.dll heap <n> MiB committed, fromsnmalloc_rs::SnMalloc::memory_stats(): snmalloc's own count of the chunks it has committed and handed to its allocators, two atomic loads, page boxes included. That is the existing pattern for a PE-side memory figure that every build reports, and it needs no new field in the per-frame payload; the walk runs only on the watch's debug samples and its threshold warnings.perf_ruleinwindows/tests/tests/e2e/bench.rsrecordsprocess_footprint_bytesandmetal_allocated_bytesasinfo: they count Wine, the driver and what earlier benchmarks of a round left in the process, so they move with the machine and the round's order.tex_staging_wrapped_byteskeeps thebytesrule, since it counts the layer's own cache, which the frames' API calls decide.docs/ARCHITECTURE.mdlists the three keys, the keys a window can leave out and the_bytessuffix's sampled form, describes the new holder and the heap figure in the address-space watch section with a reworked example line, and states the wrapper's retirement under "Retirement counters are exact".A known limit of the bound: a level the game rewrote every frame for a while and then leaves alone keeps its one kept wrapper, and so that level's staging pages, until the backing changes or the texture is destroyed, as every level did before this change. That is at most one level's staging per such texture.
Alternatives considered
Retiring the cached wrapper only once the PE side has actually dropped the staging, by carrying the drop back to the encoder in a later frame. That retires exactly, the rewritten-then-released level included, but needs a new frame operation from the API thread. Remembering the retired backing in the slot needs no wire change and bounds the churn to one extra wrapper per backing; its residual is the level that is rewritten every frame for a while and released later, which holds its pages until texture destroy as before.
Not caching wrappers for droppable levels at all. A DEFAULT-pool level written by several
UpdateTexturerects before its staging is fully written uploads more than once from the same pages, and those uploads would each create a wrapper; retiring at the release keeps the cache for them.For the heap figure, carrying snmalloc's statistics in the per-frame payload to the perf summary, as the
pe_pagebox_*keys are. That adds wire fields to every frame for a figure that changes slowly and that the address-space watch, which every build has, already sits next to.For the upload-lease holder, a process-wide gauge charged when a lease is created, as
held_pagesdoes for surfaces and encoder leases. A read lease shares its owner with the texture until the texture lets go, so such a gauge would count every live texture's staging twice; the sample-time walk counts an allocation only once nothing but leases holds it.Verification
make fmt,make checkandmake check PERF=1are green. They cover rustfmt, clippy withnurseryandpedantic, the audit (derive inventory, doc-block shape, test file placement, the COVERAGE index, the confined patterns) and the docs.make test-unitpasses (1996 core tests, 758 unix and runner tests), and so doesmake test-unit PERF=1(2065 and 760);make test-unit PERF=1also passes on the first commit alone.make test ISOLATED=1on 487f637 passes both end-to-end legs: i686 1085 tests, 1072 passed, 0 failed, 13 ignored, 6 processes; x86_64 1083 tests, 1070 passed, 0 failed, 13 ignored, 6 processes.The fix is tested at helper level.
encoder::tests::taking_a_staging_wrapper_empties_its_slot_and_keeps_the_otherspins the retirement helper, which empties one slot and hands the pages' owner over with the wrapper.encoder::tests::a_release_retires_a_levels_wrapper_once_per_backingpins the churn bound: the first answer retires and the slot remembers the backing, a wrapper recreated over that backing survives later answers, a backing change still retires it, and a never-released slot, a live wrapper or another length start over.encoder::tests::the_wrapper_gauge_counts_populated_level_slots_onlypins the wrapper gauge.guest_pages::tests::lease_only_pages_count_staging_the_texture_let_go_of_oncepins the new holder: staging a texture holds is not counted, released staging counts once for two leases, a partial tally claims nothing, and a lease whose read was released still counts. The summary andperf-kvgolden tests carry the new rows and keys;perf::tests::an_unsampled_window_prints_n_a_and_leaves_the_memory_keys_outpins a window without the memory sample (the three cells readn/ain their columns, the three keys are absent, nothing else moves); the holder tests inaddress_spacecarry the new clause.No new end-to-end test, because the suite cannot observe the fix. The pages it frees are a DEFAULT-pool level's staging, which the application never sees mapped; the gauges that count them are statics of
d3d9.dllthat a test executable cannot read; and a freed staging box is often parked in the staging lane rather than decommitted, so not even the address space changes visibly. The suite already exercises every path that runs after a released level: re-uploads after the release through partialUpdateTexture,UpdateSurface,LockRectandGetDC(see thetextures.rsrow ofwindows/tests/COVERAGE.md). After this change each of them creates a fresh wrapper instead of replacing a stale one, and all of them pass.make bench-ab BASE=5eb2ec52 BENCH_WAIT_IDLE=120(defaultwowset, five round pairs, production i686 build): 9 benchmarks, 3125 metrics, 1 regressed, 3 improved, 0 exact changed, 35 added (the five new keys), 2 incomplete (the fault keys ofstreaming, absent from one base window), and both scene shapes unchanged. The one regression isperf.enc_op_binds_vbib_msinwow112, 0.173 to 0.181 ms (+4.62 %; base rounds 0.173, 0.173, 0.174, 0.173, 0.178, candidate 0.181, 0.169, 0.181, 0.181, 0.178). That is the per-draw VB/IB bind phase, which this change does not touch, in a benchmark that uploads no textures (tex_uploads_pf0 in both legs, wrapper creates and retires 0), so none of the new code runs there apart from the once-per-window gauge read; this reads as code placement, the kind of gap an A/A run of this benchmark has shown before and the same span as #1077, andwow112'sperf.enc_ms(+0.71 %, sigma 0.83 %) andframe.p50(-0.17 %) agree. The improvements areenc_op_resolve_msand its children inapi_call_cost(-11 %).frame.p50per benchmark, base to candidate: buffers 0.217 to 0.218 ms, frame_shape 0.543 to 0.544, query_poll_spec 0.371 to 0.372, query_poll_wow 0.055 to 0.055, streaming 0.188 to 0.195 (+1.95 %, sigma 3.42 %), wow112 2.060 to 2.056. Submit and encode:submit_ms+3.16 % in wow112 (sigma 2.69 %), +0.55 % in frame_shape, +1.61 % in streaming (sigma 5.14 %);enc_work_ms+0.71 %, +3.41 % (sigma 3.33 %) and +3.12 % (sigma 2.57 %); none flagged. Instreaming, the only benchmark that uploads, every texture upload row is unchanged (9.583 uploads a frame: 7.583 raw, 1.500 padded, 0.500 pass),tex_destroy_pfis 13.082 against 13.084, and the candidate creates and retires 9.583 wrappers a frame, one per upload, since that benchmark's textures live one frame and each upload wraps a new texture's staging; the destroy count does not move because those wrappers used to retire at texture release in the same frame.GTA IV on 487f637 (
GTAIV-3489.log) against e99fa02 (GTAIV-12703.log) at matched session times: page-boxotherstays at 0 MiB for the whole run (before 107 to 617 MiB), the page boxes stay flat at 425 to 445 MiB (before 523 to 1052 MiB), and at +4:34 the real free address space reads 1801 against 1238 MiB with the largest free block at 1758 against 1167 MiB.upload leasespeaks at 6 MiB andwrappedat 7.8 MiB. Wrapper churn over the run is 56287 created and 56187 retired, a median of 1.35 a frame with spikes to 30 a frame while the game streams. The process footprint still grows by 237 MiB from 02:40:36 to the end of the run; of that, the Metal allocated size accounts for 3 MiB,d3d9.dll's heap for 8 and the page boxes for 1, and the rest is the game's and Wine's own 32-bit memory (about 133 MiB) and the 64-bit side (about 94 MiB), outside what this layer allocates. The second half of that run was submit-bound (submit 8.74 against 6.6 ms, 109 against 120 fps) with no correlation to the wrapper churn in the same windows; it is likely the scene, and it is not explained here.Rules check
Statics: none added. The heap figure reads snmalloc's process-wide backend counters;
LeaseOnlyPagesis a value built per sample.Config keys and environment variables: none.
Wire fields: none. The fix uses
release_staging, au32already onTextureUploadRecord. The new counters are read on the side that reports them: the footprint, the Metal size and the wrapper bytes on the unix side, the heap and the upload leases on the PE side.perf-kvkeys: five added,process_footprint_bytes,metal_allocated_bytes,tex_staging_wrapped_bytes,tex_wrapper_create_totalandtex_wrapper_retire_total, each with its row in the ARCHITECTURE table and its bench class. No existing key changes name, unit or aggregation.Dependencies: none.
proc_pid_rusageandrusage_info_v2come fromlibc,memory_statsfrom the pinnedsnmalloc-rs.Derives:
LeaseOnlyPagesderivesDefaultonly; noCloneorCopywas added, so the derive inventory is unchanged.MipStagingBuffer, which already derivesCloneandDefault, gains oneboolfield, its only one, so the booleans-pack rule does not apply.Lint suppressions: none. Unsafe: two blocks in
process_footprint, one operation each with its SAFETY comment.Duplicates:
park_staging_wrapperreplaces the inline retirement at the backing change and at texture destroy, so all three retirement sites share it and its counter.process_footprintsits besidetask_faultsinhandlers.rsand follows it.Placement: the lease tally is pure and lives in
mtld3d-corewith its unit test; the wrapper helpers are free functions tested inunix/unix/src/encoder/tests.rs; every test is in its<stem>/tests.rs. The summary rows go throughres_rowand the golden layout and kv tests carry them.