Conversation
jwalk's parallel mode reads directories on the rayon pool as fast as it
can and holds each one's entries until the iterator reaches them. The
bounded channel only ever held back the thread feeding it, so on a
100,000-file library about 25 MB of directory entries — nearly the whole
tree — were in memory at the scan's peak. Walking serially keeps one
directory at a time; the stat stage behind it is already concurrent, so
the scan is no slower (warm 0.8 s either way, cold 14 s before and 9 s
after, now that tag reading no longer competes with eight walkers).
The fingerprint map is the other allocation that grows with the library.
Its values now hold whole seconds, which is all the index stores, rather
than two 16-byte SystemTimes, and its keys are boxed paths relative to
the root being scanned. The table at 100,000 files goes from 10.1 MB to
7.1 MB and each key is shorter by the root. The two batch vectors are
no longer reserved up front, which held 1.3 MB through every scan of an
unchanged library.
Measured with a counting allocator on a 100,000-file library:
cold scan peak heap 36.3 MB -> 11.3 MB
warm scan peak heap 58-61 MB -> 41 MB
max RSS (warm) 107 MB -> 81 MB
idle footprint after 65 MB -> 39-40 MB (malloc fragmentation
left by the peak: 35 MB -> 13 MB)
threads at idle 22 -> 14 (no rayon pool until a transcode)
libxaac sizes an encoder for USAC whatever the profile: its state carries
a complete USAC encoder by value and its API struct a USAC configuration
per element, 44 MB and 12 MB. An AAC-LC encoder touches well under a
megabyte of that, but the library clears every block it is handed, so all
55 MB became resident for as long as the encoder lived — for each browser
fMP4 or HLS stream, each audio.aac request, whatever the channel count.
xaac-rs already hands the encoder zeroed memory, so the clears are
redundant. libxaac-sys now builds the encoder with IXHEAACE_ZEROED_ALLOC
(through a LIBXAAC_ZEROED_ALLOC cmake option, as LIBXAAC_ARM_FLOAT_ABI is
passed), which skips the five that cover freshly allocated blocks. That
alone only helped the first encoder: the heap keeps freed blocks this
size and clears a reused one to honour alloc_zeroed, which writes every
page again. So on Unix xaac-rs maps blocks of a megabyte or more directly
and unmaps them on drop; smaller blocks and other platforms keep the heap.
Output is byte-identical to the unpatched encoder (20 s of stereo and of
5.1, compared with cmp). Measured:
one encoder, RSS after create +55 MB -> +0 MB
create/encode/drop, repeated +60 MB each -> ~1 MB
server footprint after one
audio.aac transcode 75 MB -> 19 MB
DTS film as browser fMP4 86 MB -> 32 MB
6 concurrent AAC encodes 388 MB -> 48 MB (max RSS 413 -> 65)
The fingerprint map and the root its keys are relative to travelled as two arguments, which took read_window past clippy's argument limit and left every caller translating paths by hand. IndexedTree owns both: lookups take the walk's absolute paths, and deletions come back out absolute.
The scan built its compact map from load_file_fingerprints_under, which returned the whole subtree at once: every record with its absolute path, in a vector that grew by doubling. At 100,000 files that was 18 MB of paths and a 10 MB vector alive beside the map being built from them, and it had become the scan's peak. The loader now takes a cursor and a limit and returns one page in path order, a keyset query that seeks the path index to where the previous page stopped. The scanner converts each 4,096-row page (about a megabyte) into the map before reading the next, so the load costs a page on top of the map instead of the subtree twice over. A cursor below the range is ignored rather than allowed to widen it.
Both the macOS allocator and glibc keep freed pages mapped for reuse, so whatever a scan peaked at stayed counted against the process for the rest of its run: after a 100,000-file startup scan, with the scan's own peak already reduced, 30 MB of malloc memory was still dirty for 17 MB in use. Once a scan that covered at least 10,000 files has dropped its index, the scanner asks the allocator to release what is free (malloc_zone_pressure_relief on macOS, malloc_trim on glibc; nothing elsewhere), on the blocking pool. A watcher's rescan of one folder is below the threshold and never pays for the call.
Three caches in TranscodeState could together hold about 136 MB for the
life of the process, and every entry past the first few was insurance
against a request that rarely comes:
segments 48 MB -> 24 MB what is asked for twice is a segment just
played (a seek back, a re-buffer), so only the
recent end of each stream earns its place
seek runs 64 MB -> 24 MB eviction is oldest-first and a television's
repeats come at the end of its binary search,
once the probes have narrowed to one group of
pictures; those last runs are what is kept
indexes 8 -> 4 ~3 MB each for a two-hour track, needed across
the few requests that open one file
The worst case falls from about 136 MB to 60 MB. Entry counts are
unchanged, and none of this touches a stream that is playing.
The cache pairs the halves of a rename on macOS and Windows and holds up to two generations of 25,000 paths. Every entry was an absolute PathBuf, so each one repeated its root's prefix: in a profile of a 100,000-file library it was the largest thing in the idle heap, 6.4 MB for one generation. Entries are now kept per watch root, keyed by a boxed path below it, which saves the root's length plus eight bytes per entry — 2.4 MB at the cap for a 40-byte root. A directory the watcher sees created inside a root is seeded into that root rather than registered as another, so the set of roots cannot grow with the library. Lookups still allocate nothing; removal drops a subtree from the root that holds it, or the whole root when that is what was removed.
17d1309 capped symphonia probes at two for the whole process, to bound what the parsers allocate: symphonia declares a limit on picture bytes but no reader enforces it. The scanner, though, spreads its reads across every core on purpose (read_window's buffer_unordered), and two slots undid that. With 200 files already cached, two slots probed them in 7.9 ms and eight in 4.1 ms on this eight-core machine; on a network share, where a probe mostly waits on round trips, the gap is larger. The slots now follow the core count, at least two and at most eight, still shared by every scan and watch event, so a many-core server does not multiply the parser's peak by its core count.
17d1309 put passthrough segments behind the same four build slots as decoded ones, which bounds this path's memory, but a request that found them busy was answered 503 at once. A build keeps its slot until it finishes even when its request is gone, so a player that scrubs met the builds its own seeks had abandoned as a run of 503s, and passthrough had never been refused before. A request now waits up to five seconds for a slot, inside a player's own first-byte timeout (hls.js allows ten), and is refused only after that.
Since 17d1309 a station's queue holds ids rather than paths. A row that is removed and indexed again, as a deleted-and-restored file or an editor's write-and-rename save leaves it, comes back under a new id, and the pass under way skipped the track; loop and shuffle stations picked it up only at the next pass, and a linear station, whose single pass is its broadcast, never played it. A queue of paths had played it anyway. A lookup that misses now reads the queue again and carries on from the latest track already tried, leaving out everything the pass attempted. One re-read per run of misses, so a folder being deleted costs one rebuild rather than one per track. The new test fails on the old loop, which broadcast a and c but not b.
The audit's fragment row came from the multi-track builder, which only films with re-encoded audio go through; every HLS segment uses the single-track one, which kept one extra payload-sized copy rather than two. The RSS probe now measures either shape (VUIO_FRAGMENT_SHAPE=single): for 64 MiB of payload, 201.8 to 137.6 MiB single-track (32%) and 265.8 to 137.8 MiB multi-track (48%), all with identical output. The audit also describes the probe and segment-slot changes, and notes that every figure is a macOS peak: glibc keeps freed memory in arenas, and the Docker image is built on musl, where release_free_memory does nothing.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
A memory pass over the whole server, driven by profiling rather than guessing: the release build run under a counting allocator and dhat, with
footprint/vmmapfor what the heap numbers miss, against a 100,000-file library, AC-3 and DTS films, and multichannel audio. Seven fixes, one per allocation that was either the largest of its kind or grew with the library.audio.aacstreamPathBufper entry, 6.4 MB per 25k-entry generation1. The AAC encoder
libxaac sizes every encoder for USAC: its state carries a whole USAC encoder by value and the API struct a USAC config per element, 44 MB + 12 MB. AAC-LC touches under a megabyte, but the library cleared every block it was handed, so all 55 MB became resident.
libxaac-sysbuilds the encoder withIXHEAACE_ZEROED_ALLOC(aLIBXAAC_ZEROED_ALLOCcmake option, passed the wayLIBXAAC_ARM_FLOAT_ABIalready is), which skips the five clears of freshly allocated blocks. xaac-rs already allocated them zeroed.alloc_zeroed. So on Unix xaac-rsmmaps blocks of a megabyte or more andmunmaps them on drop. Smaller blocks and Windows keep the heap.Output is byte-identical to the unpatched encoder (20 s of stereo and of 5.1, compared with
cmp).2–5. Library scans
jwalk::Parallelism::Serial. The bounded channel only held back the thread feeding it; the rayon walkers read ahead with no bound. The scan is no slower, because every path isstated afterwards and that stage is concurrent.SystemTimes; keys areBox<Path>relative to the root.IndexedTreeowns the map and its root.load_file_fingerprints_undernow takes a cursor and a limit: a keyset query on the path index. It seeks to where the previous page stopped and ignores a cursor below the range. Thedatabasemodule is not inapi-surface.txt, so this is not a public API change.platform::release_free_memory():malloc_zone_pressure_reliefon macOS,malloc_trimon glibc, a no-op elsewhere. It runs on the blocking pool after any scan that covered at least 10,000 files, so a watcher's single-folder rescan never pays for it.6–7. Caches
Measurements
Release build on an M1, 100,000-file library, counting allocator plus
footprint. Numbers measured for fixes 1–3:audio.aac(5.1, 10 min)What bounds the warm-scan peak of 41 MB is the 28 MB load that fix 4 removes. Fixes 4–7 were not re-profiled; their numbers above are the sizes they remove, from the same profile.
Tests
fingerprints_under_a_root_come_back_in_pages_that_tile_it: pages of 3 tile an 8-record subtree exactly, in order ([3, 3, 2]). Siblings either side of the range (Film.mkv,Film0.mkv,Films/) stay out, and a cursor below the range doesn't widen it.a_root_holds_its_paths_and_the_folders_created_under_it: a folder created inside a root is found through that root and doesn't become a new one. Removing a subfolder drops exactly its four entries, and removing the root empties the cache.Both were confirmed to fail against a regression: without the cursor guard, and with every seeded directory registered as a root.
Local verification:
cargo test -p vuio-corewith every feature exceptweb-ui: all passing (654 lib tests plus every integration binary exceptweb_ui_integration_tests)cargo test -p xaac-rs: passingcargo fmt --all -- --check: cleancargo clippy -D warningsonvuio-core(lib and those test targets),xaac-rsandlibxaac-sys: cleanweb-uicouldn't be built locally, because the checked-incrates/vuio-web/distis missing from my working tree; CI covers it. The Linux (malloc_trim), FreeBSD and Windows paths are compiled only by CI.