Skip to content

Lists: one allocation per list again; [each acc, x] appends in place - #228

Merged
revarbat merged 2 commits into
mainfrom
perf/list-one-allocation
Oct 2, 2026
Merged

revarbat merged 2 commits into
mainfrom
perf/list-one-allocation

Conversation

@revarbat

@revarbat revarbat commented Oct 2, 2026

Copy link
Copy Markdown
Member

Refs BelfrySCAD/BelfrySCAD#668

This follows up 1.35.0's shared-prefix lists (#225). That change made accumulator loops linear but, measured properly, cost ordinary BOSL2 code 1–5.6% in instructions and inflated a long docs build's footprint. This PR keeps the linear appends and removes both costs. It also extends in-place appending to [each acc, x].

Why 1.35.0 regressed (profiled with sample; heap inspected with heap/vmmap):

  • One extra allocation per list. A ValueList held a shared_ptr to a separately allocated ListBuffer, which held the std::vector. That showed up as about +2.5% in malloc/free and +0.7% tearing down the ListBuffer control block.
  • ×2 growth with at least 16 slots produced large freed blocks. macOS keeps those resident: after the shapes3d docs the live heap was unchanged (74.6 vs 74.4 MB), but 305 MB of freed large blocks were still held.

Changes:

  • makeList: the list and its buffer header share one block, and the element vector is adopted, not copied. That's the same two allocations a plain vector cost.
  • Cached element pointer: each list caches it, so a read is one hop.
  • Appended lists: a list appended in place is a small ValueList viewing its parent's buffer, keeping the parent's block alive through a shared_ptr that points into it.
  • Construction: ListItems is neither copyable nor movable, so all list construction goes through makeList, which the compiler enforces.
  • Growth is ×1.5 with no minimum.
  • [each acc, x]: a list literal or comprehension whose first contribution is each <list> now starts as that list (ListBuilder), on the VM and the interpreter. It was still quadratic: 50k steps took 6.1 s; 400k now take 0.2 s.

Measured with instructions retired (/usr/bin/time -l), which load doesn't affect. Wall-clock timings on this machine were distorted by unrelated load, and that's how #225 shipped a false claim.

Nine BOSL2 scripts, against 1.34.0:

script 1.35.0 this PR
paths +5.6% +0.8%
vnf +3.5% +0.2%
skin +2.6% +0.5%
rounding +2.5% +0.1%
cuboids +1.8% +0.5%
threads +2.2% −0.9%
sweep +1.3% −0.4%
text +1.3% +0.9%
gears +1.0% +0.1%

Bytes allocated: 1.35.0 was +1.4% to +33.5% over 1.34.0; this PR is +1.6% to +24%.

Full BOSL2 docs build:

instructions cycles peak footprint
1.34.0 7,522.6 G 1,862.2 G 1,831 MB
1.35.0 7,753.8 G (+3.1%) 1,915.9 G 1,784 MB
this PR 7,526.2 G (+0.05%) 1,861.9 G 1,651 MB

shapes3d docs, after the run:

  • 1.34.0: footprint 380 MB, 3.6 MB of freed large blocks held.
  • 1.35.0: 738 MB footprint, 305 MB held.
  • This PR: 390 MB footprint, 3.6 MB held.

Append loops are still linear: a million concat appends take 0.6 s, and #668's script stops with "Recursion detected" in under 2 s.

Tested:

  • ctest: 2014 passed.
  • AddressSanitizer: all 1355 tests, plus the append, #668 and BOSL2 scripts, clean.
  • macOS leaks: 0.
  • test_python_bindings.py: 61 passed, on a wheel from this branch.
  • BelfrySCAD suite: 2193 passed.
  • BOSL2 docs A/B against 1.35.0: 56/56 markdown files identical and logs identical. 2636/2640 images are pixel-identical. The other 4 are cubetruss images with 1–28 edge pixels changed. Those come from Patch Manifold v3.5.2: fix a heap overflow in CsgLeafNode::Compose #227's Manifold patch, not from this PR: an unpatched build of this branch gives cubetruss meshes identical to 1.35.0's.

🤖 Generated with Claude Code

revarbat and others added 2 commits October 2, 2026 03:58
The other accumulator idiom, f(i, [each acc, x]), copied acc on every
step just as concat used to. A list literal or comprehension whose first
contribution is `each <a list>` now starts as that list and appends the
rest with listAppend (ListBuilder), on the VM's Accum* ops and the
interpreter's evalListLiteral alike. A string, range or undef after
`each` still expands as before, and `each` anywhere but first is
unchanged.

50,000 steps of [each acc, i]: 6.1 s on 1.35.0, quadratic; 100,000 now
take 0.06 s, and 400,000 0.19 s.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Shared-prefix lists (1.35.0) gave each list a ValueList plus a separate
shared_ptr-owned ListBuffer holding the vector: one more allocation per
list than a plain vector, and more and odd-sized freed memory. Measured
by instructions retired against 1.34.0, ordinary BOSL2 ran 1-5.6% slower,
and a long docsgen's footprint rose 63% from freed large blocks the
allocator kept (live heap unchanged).

Now makeList puts the list and its buffer header in one block, adopting
the element vector rather than copying it -- the same two allocations a
plain vector cost -- and each list caches its element pointer, so a read
is one hop. A list appended in place is a small ValueList viewing its
parent's buffer and keeping the parent's block alive. Growth is x1.5 with
no minimum, instead of x2 with at least 16 slots. ListItems is neither
copyable nor movable, since an own-buffer view must not outlive its block;
all list construction goes through makeList.

Instructions retired vs 1.34.0 on nine BOSL2 scripts: -0.9% to +0.9%
(1.35.0: +1.0% to +5.6%). Bytes allocated: +1.6% to +24% (1.35.0: +1.4%
to +33.5%). A million appends still take 0.6 s.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@revarbat
revarbat merged commit 74fdf7a into main Oct 2, 2026
3 checks passed
@revarbat
revarbat deleted the perf/list-one-allocation branch October 2, 2026 11:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant