Skip to content

Cut the server's largest memory allocations - #55

Open
vyrti wants to merge 12 commits into
mainfrom
perf/memory-allocations
Open

vyrti wants to merge 12 commits into
mainfrom
perf/memory-allocations

Conversation

@vyrti

@vyrti vyrti commented Sep 23, 2026

Copy link
Copy Markdown
Contributor

Summary

A memory pass over the whole server, driven by profiling rather than guessing: the release build run under a counting allocator and dhat, with footprint/vmmap for what the heap numbers miss, against a 100,000-file library, AC-3 and DTS films, and multichannel audio. Seven fixes, one per allocation that was either the largest of its kind or grew with the library.

# Allocation Before After
1 AAC encoder (libxaac), one per browser fMP4 / HLS / audio.aac stream 55 MB resident each, any channel count under 1 MB each
2 jwalk reading the tree ahead of the scan ~25 MB at 100k files, grows with the library one directory at a time
3 The scan's fingerprint map 10.1 MB table + absolute keys at 100k 7.1 MB table + root-relative keys
4 The subtree's rows loaded at scan start 18 MB of paths + 10 MB vector at 100k one 4,096-row page (~1 MB)
5 Freed memory the allocator keeps after a scan 30 MB dirty for 17 MB in use released after scans of 10k+ files
6 Transcode caches (segments / seek runs / indexes) up to ~136 MB up to ~60 MB
7 Watcher file-id cache absolute PathBuf per entry, 6.4 MB per 25k-entry generation keyed per root, root prefix + 8 B saved per entry

1. The AAC encoder

libxaac sizes every encoder for USAC: its state carries a whole USAC encoder by value and the API struct a USAC config per element, 44 MB + 12 MB. AAC-LC touches under a megabyte, but the library cleared every block it was handed, so all 55 MB became resident.

  • libxaac-sys builds the encoder with IXHEAACE_ZEROED_ALLOC (a LIBXAAC_ZEROED_ALLOC cmake option, passed the way LIBXAAC_ARM_FLOAT_ABI already is), which skips the five clears of freshly allocated blocks. xaac-rs already allocated them zeroed.
  • That alone helped only the first encoder: the heap caches freed blocks this size and has to clear a reused one to honour alloc_zeroed. So on Unix xaac-rs mmaps blocks of a megabyte or more and munmaps them on drop. Smaller blocks and Windows keep the heap.

Output is byte-identical to the unpatched encoder (20 s of stereo and of 5.1, compared with cmp).

2–5. Library scans

  • 2 jwalk::Parallelism::Serial. The bounded channel only held back the thread feeding it; the rayon walkers read ahead with no bound. The scan is no slower, because every path is stated afterwards and that stage is concurrent.
  • 3 Map values hold whole seconds (all the index stores) instead of two SystemTimes; keys are Box<Path> relative to the root. IndexedTree owns the map and its root.
  • 4 load_file_fingerprints_under now takes a cursor and a limit: a keyset query on the path index. It seeks to where the previous page stopped and ignores a cursor below the range. The database module is not in api-surface.txt, so this is not a public API change.
  • 5 platform::release_free_memory(): malloc_zone_pressure_relief on macOS, malloc_trim on glibc, a no-op elsewhere. It runs on the blocking pool after any scan that covered at least 10,000 files, so a watcher's single-folder rescan never pays for it.

6–7. Caches

  • 6 Segments 48 → 24 MB, seek runs 64 → 24 MB, indexes 8 → 4. Eviction is oldest-first, and what gets asked for again is the recent end: a segment just played, or the last runs of a television's binary search once it has narrowed to one group of pictures.
  • 7 The file-id cache (rename pairing on macOS/Windows) stores entries per watch root. A directory created inside a root is seeded into that root rather than becoming another one, so the set of roots can't grow with the library.

Measurements

Release build on an M1, 100,000-file library, counting allocator plus footprint. Numbers measured for fixes 1–3:

before after
cold scan, peak heap 36.3 MB 11.3 MB
cold scan, time 14 s 9 s
warm scan (all unchanged), peak heap 58–61 MB 41 MB
warm scan, max RSS 107 MB 81 MB
idle footprint after the startup scan 65 MB 39–40 MB
threads at idle 22 14
footprint after one audio.aac (5.1, 10 min) 75 MB 19 MB
footprint after a DTS film as browser fMP4 86 MB 32 MB
footprint, 6 concurrent AAC encodes 388 MB 48 MB

What bounds the warm-scan peak of 41 MB is the 28 MB load that fix 4 removes. Fixes 4–7 were not re-profiled; their numbers above are the sizes they remove, from the same profile.

Tests

  • fingerprints_under_a_root_come_back_in_pages_that_tile_it: pages of 3 tile an 8-record subtree exactly, in order ([3, 3, 2]). Siblings either side of the range (Film.mkv, Film0.mkv, Films/) stay out, and a cursor below the range doesn't widen it.
  • a_root_holds_its_paths_and_the_folders_created_under_it: a folder created inside a root is found through that root and doesn't become a new one. Removing a subfolder drops exactly its four entries, and removing the root empties the cache.

Both were confirmed to fail against a regression: without the cursor guard, and with every seeded directory registered as a root.

Local verification:

  • cargo test -p vuio-core with every feature except web-ui: all passing (654 lib tests plus every integration binary except web_ui_integration_tests)
  • cargo test -p xaac-rs: passing
  • cargo fmt --all -- --check: clean
  • cargo clippy -D warnings on vuio-core (lib and those test targets), xaac-rs and libxaac-sys: clean
  • web-ui couldn't be built locally, because the checked-in crates/vuio-web/dist is missing from my working tree; CI covers it. The Linux (malloc_trim), FreeBSD and Windows paths are compiled only by CI.

jwalk's parallel mode reads directories on the rayon pool as fast as it
can and holds each one's entries until the iterator reaches them. The
bounded channel only ever held back the thread feeding it, so on a
100,000-file library about 25 MB of directory entries — nearly the whole
tree — were in memory at the scan's peak. Walking serially keeps one
directory at a time; the stat stage behind it is already concurrent, so
the scan is no slower (warm 0.8 s either way, cold 14 s before and 9 s
after, now that tag reading no longer competes with eight walkers).

The fingerprint map is the other allocation that grows with the library.
Its values now hold whole seconds, which is all the index stores, rather
than two 16-byte SystemTimes, and its keys are boxed paths relative to
the root being scanned. The table at 100,000 files goes from 10.1 MB to
7.1 MB and each key is shorter by the root. The two batch vectors are
no longer reserved up front, which held 1.3 MB through every scan of an
unchanged library.

Measured with a counting allocator on a 100,000-file library:

  cold scan peak heap     36.3 MB -> 11.3 MB
  warm scan peak heap     58-61 MB -> 41 MB
  max RSS (warm)          107 MB -> 81 MB
  idle footprint after    65 MB -> 39-40 MB (malloc fragmentation
                          left by the peak: 35 MB -> 13 MB)
  threads at idle         22 -> 14 (no rayon pool until a transcode)
libxaac sizes an encoder for USAC whatever the profile: its state carries
a complete USAC encoder by value and its API struct a USAC configuration
per element, 44 MB and 12 MB. An AAC-LC encoder touches well under a
megabyte of that, but the library clears every block it is handed, so all
55 MB became resident for as long as the encoder lived — for each browser
fMP4 or HLS stream, each audio.aac request, whatever the channel count.

xaac-rs already hands the encoder zeroed memory, so the clears are
redundant. libxaac-sys now builds the encoder with IXHEAACE_ZEROED_ALLOC
(through a LIBXAAC_ZEROED_ALLOC cmake option, as LIBXAAC_ARM_FLOAT_ABI is
passed), which skips the five that cover freshly allocated blocks. That
alone only helped the first encoder: the heap keeps freed blocks this
size and clears a reused one to honour alloc_zeroed, which writes every
page again. So on Unix xaac-rs maps blocks of a megabyte or more directly
and unmaps them on drop; smaller blocks and other platforms keep the heap.

Output is byte-identical to the unpatched encoder (20 s of stereo and of
5.1, compared with cmp). Measured:

  one encoder, RSS after create      +55 MB -> +0 MB
  create/encode/drop, repeated       +60 MB each -> ~1 MB
  server footprint after one
    audio.aac transcode              75 MB -> 19 MB
    DTS film as browser fMP4         86 MB -> 32 MB
    6 concurrent AAC encodes         388 MB -> 48 MB (max RSS 413 -> 65)
The fingerprint map and the root its keys are relative to travelled as two
arguments, which took read_window past clippy's argument limit and left
every caller translating paths by hand. IndexedTree owns both: lookups
take the walk's absolute paths, and deletions come back out absolute.
The scan built its compact map from load_file_fingerprints_under, which
returned the whole subtree at once: every record with its absolute path,
in a vector that grew by doubling. At 100,000 files that was 18 MB of
paths and a 10 MB vector alive beside the map being built from them, and
it had become the scan's peak.

The loader now takes a cursor and a limit and returns one page in path
order, a keyset query that seeks the path index to where the previous
page stopped. The scanner converts each 4,096-row page (about a megabyte)
into the map before reading the next, so the load costs a page on top of
the map instead of the subtree twice over. A cursor below the range is
ignored rather than allowed to widen it.
Both the macOS allocator and glibc keep freed pages mapped for reuse, so
whatever a scan peaked at stayed counted against the process for the
rest of its run: after a 100,000-file startup scan, with the scan's own
peak already reduced, 30 MB of malloc memory was still dirty for 17 MB
in use. Once a scan that covered at least 10,000 files has dropped its
index, the scanner asks the allocator to release what is free
(malloc_zone_pressure_relief on macOS, malloc_trim on glibc; nothing
elsewhere), on the blocking pool. A watcher's rescan of one folder is
below the threshold and never pays for the call.
Three caches in TranscodeState could together hold about 136 MB for the
life of the process, and every entry past the first few was insurance
against a request that rarely comes:

  segments  48 MB -> 24 MB  what is asked for twice is a segment just
                            played (a seek back, a re-buffer), so only the
                            recent end of each stream earns its place
  seek runs 64 MB -> 24 MB  eviction is oldest-first and a television's
                            repeats come at the end of its binary search,
                            once the probes have narrowed to one group of
                            pictures; those last runs are what is kept
  indexes   8 -> 4          ~3 MB each for a two-hour track, needed across
                            the few requests that open one file

The worst case falls from about 136 MB to 60 MB. Entry counts are
unchanged, and none of this touches a stream that is playing.
The cache pairs the halves of a rename on macOS and Windows and holds up
to two generations of 25,000 paths. Every entry was an absolute PathBuf,
so each one repeated its root's prefix: in a profile of a 100,000-file
library it was the largest thing in the idle heap, 6.4 MB for one
generation.

Entries are now kept per watch root, keyed by a boxed path below it,
which saves the root's length plus eight bytes per entry — 2.4 MB at the
cap for a 40-byte root. A directory the watcher sees created inside a
root is seeded into that root rather than registered as another, so the
set of roots cannot grow with the library. Lookups still allocate
nothing; removal drops a subtree from the root that holds it, or the
whole root when that is what was removed.
17d1309 capped symphonia probes at two for the whole process, to bound
what the parsers allocate: symphonia declares a limit on picture bytes
but no reader enforces it. The scanner, though, spreads its reads across
every core on purpose (read_window's buffer_unordered), and two slots
undid that. With 200 files already cached, two slots probed them in
7.9 ms and eight in 4.1 ms on this eight-core machine; on a network
share, where a probe mostly waits on round trips, the gap is larger.

The slots now follow the core count, at least two and at most eight,
still shared by every scan and watch event, so a many-core server does
not multiply the parser's peak by its core count.
17d1309 put passthrough segments behind the same four build slots as
decoded ones, which bounds this path's memory, but a request that found
them busy was answered 503 at once. A build keeps its slot until it
finishes even when its request is gone, so a player that scrubs met the
builds its own seeks had abandoned as a run of 503s, and passthrough had
never been refused before.

A request now waits up to five seconds for a slot, inside a player's own
first-byte timeout (hls.js allows ten), and is refused only after that.
Since 17d1309 a station's queue holds ids rather than paths. A row that
is removed and indexed again, as a deleted-and-restored file or an
editor's write-and-rename save leaves it, comes back under a new id, and
the pass under way skipped the track; loop and shuffle stations picked
it up only at the next pass, and a linear station, whose single pass is
its broadcast, never played it. A queue of paths had played it anyway.

A lookup that misses now reads the queue again and carries on from the
latest track already tried, leaving out everything the pass attempted.
One re-read per run of misses, so a folder being deleted costs one
rebuild rather than one per track. The new test fails on the old loop,
which broadcast a and c but not b.
The audit's fragment row came from the multi-track builder, which only
films with re-encoded audio go through; every HLS segment uses the
single-track one, which kept one extra payload-sized copy rather than
two. The RSS probe now measures either shape (VUIO_FRAGMENT_SHAPE=single):
for 64 MiB of payload, 201.8 to 137.6 MiB single-track (32%) and 265.8 to
137.8 MiB multi-track (48%), all with identical output.

The audit also describes the probe and segment-slot changes, and notes
that every figure is a macOS peak: glibc keeps freed memory in arenas,
and the Docker image is built on musl, where release_free_memory does
nothing.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant