Skip to content

Share list storage between prefixes so concat appends in place - #225

Merged
revarbat merged 1 commit into
mainfrom
perf/shared-prefix-lists
Oct 2, 2026
Merged

revarbat merged 1 commit into
mainfrom
perf/shared-prefix-lists

Conversation

@revarbat

@revarbat revarbat commented Oct 2, 2026 •

Copy link
Copy Markdown
Member

Correction (added after merge). The BOSL2 timings below were wall-clock runs on a machine with heavy unrelated load, and they were wrong.

Measured instead by instructions retired, which doesn't depend on load, against 1.34.0, ordinary BOSL2 scripts run 1–6% slower on 1.35.0:

  • paths +5.6%, vnf +3.5%, skin +2.6%, rounding +2.5%, cuboids +1.8%, gears +1.0%;
  • they also allocate 2.5–33% more bytes.

A full BOSL2 docs build used 773 s of CPU against 737 s, and its peak memory rose from 1.67 GB to 2.72 GB.

The quadratic-to-linear win for growing-list loops (and BelfrySCAD#668) is real. The likely causes of the slowdown:

  • an extra pointer hop on every element access (list → buffer → vector data);
  • at least 16 slots reserved when a list that came from appending is copied.

A rework that keeps the linear appends without either cost is in progress.

Refs BelfrySCAD/BelfrySCAD#668

Problem. concat copied its whole first argument every time. That made the idiomatic tail-recursive accumulator f(i, concat(acc, [x])) quadratic: 40k appends took 3.6 s, and a million would take an estimated ~40 minutes.

BelfrySCAD#668's script loops forever by mistake. It looked like recursion detection was broken, but the million-iteration tail-call cap does fire; it would just take far longer than anyone waits to get there. OpenSCAD avoids this with lazy concatenation: concat embeds its argument vectors and flattens them on first index.

Change. A list is now a view of the first n elements of a buffer that longer lists may share (ListItems, ListBuffer, listAppend in value.hpp/value.cpp).

  • In-place extend: concat(acc, …) extends acc's buffer in place when acc ends where the buffer's used part ends and there is spare capacity. One compare-and-swap claims the new slots, so two lists sharing a prefix can never write the same slot; the loser copies.
  • Copy sizing: a list that came from appending before is copied into double the room. Anything else is copied at its exact size, so a one-off concat(a, b) costs no spare memory.
  • Shared copies stay valid: older lists keep seeing only their own prefix, whatever still holds them. This also works through the interpreter's retained tail-call chain, so memory under the debugger is now linear too (it was quadratic).
  • Reader interface: ListItems keeps the old vector's read interface (size, [], iterators, data). Its conversion to std::vector is explicit, so a copy can never be made by accident binding a const std::vector&. Every such site was found by the compiler and moved to const ListItems&.

Results (Release build with TBB; baseline is main built with the same options):

before after
1M appends, then len ~40 min (est.) 0.48 s
growing-list loop to the 1M recursion cap ~40 min (est.) 0.44 s, correct "Recursion detected"
append and index the accumulator every step, 200k quadratic 0.13 s (lazy concat can't fix this case)
#668's script runs indefinitely "Recursion detected calling function '_getbreaks'" in 1.5 s (OpenSCAD: 5 s)
interpreter path (OSCAD_BYTECODE_VM=0), memory quadratic linear (200k appends: 231 MB)

BOSL2 timings, best of 5:

  • Faster: resample_path/path_cut_points/offset −9%, skin −8%, the reporter's text-wrapping library −8%, threaded_rod −15%.
  • Slightly slower: spur_gear and rounded cuboid +1–2% when runs are interleaved. That's the extra allocation per list (list header, shared buffer header, element storage); storing elements inline with the buffer would remove it.

Tested

  • ctest: 2013 passed, including 5 new ListAppend tests:
    • the base list is unchanged;
    • copies are logarithmic over 10k appends;
    • two branches from one prefix stay apart;
    • a one-off concat is exactly sized;
    • empty and null edge cases.
  • tests/test_python_bindings.py: 61 passed, against a wheel built from this branch. I confirmed the new code was loaded by timing 1M appends through BelfrySCAD (1.0 s).
  • BelfrySCAD's suite on this branch: 2193 passed.
  • Full BOSL2 docsgen A/B against 1.34.0: 56/56 markdown files identical, 2640/2640 images pixel-identical, logs identical. The build took 665 s against 830 s.

Not done here:

  • The [each acc, x] append form still copies. It goes through the comprehension accumulator, not concat.
  • Inline element storage, which would remove the remaining per-list allocation.

🤖 Generated with Claude Code

A list is now a view of the first n elements of a buffer that longer
lists may share. concat(acc, ...) extends acc's buffer in place when acc
ends at the buffer's frontier and there is room (claiming the slots with
one compare-and-swap), and otherwise copies -- into double the room if
the list came from appending before, exactly sized if not. Old lists
stay valid whatever still holds them, so this works through the
interpreter's retained call chain too.

The idiomatic tail-recursive accumulator, f(i, concat(acc, [x])), goes
from quadratic to linear: a million appends in 0.44 s against an
estimated 40 minutes, and the script in BelfrySCAD#668 now reaches the
recursion limit in 1.5 s. ListItems keeps the old vector's read
interface; its conversion to std::vector is explicit so no copy happens
by accident.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@revarbat
revarbat merged commit b064399 into main Oct 2, 2026
3 checks passed
@revarbat
revarbat deleted the perf/shared-prefix-lists branch October 2, 2026 04:07
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant