Audit cleanup - #4
Merged
Merged
Conversation
…shorthand classes) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
bench_compare_pcre2.py carried copies of print_comparison/speedup_bar and gen_pdf's build_table_data/generate_pdf differing only in labels, widths and titles. Parametrize the originals (defaults unchanged) and import them. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…-char forwarders PikeVM._check_anchor/_is_word_char duplicated _bt_check_anchor; the _bt_/_sbt_is_word_char forwarders become is_word_byte, and the WORD / NOT_WORD arms share one body. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A one-off cross-check CI never runs (--check-baseline refused it). Removes parse_gcov_intermediate, the gcov branches in cover_one, the flag and its guard, the opt/llvm-cov lookup and TestGcovParse. compiler-rt stays: the default link pulls libclang_rt.osx.a (clang -### shows it), and clang_impl depends on it anyway; its pixi.toml comment now says so. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…them all) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Regex[pattern, flags: RegexFlags = RegexFlags()] compiles the flags as a leading inline group (after any (*UTF8) verbs) through Regex._pat, which is exactly `pattern` for the default flags, so default instantiations keep the same memoized NFA and backtracker symbols. Drops the unused RegexFlags.__and__ and fixes the README Flags section. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Replaces _sbt_to_lower, _bt_to_lower and nfa.mojo's _to_lower (same ASCII A-Z rule; generic over the scalar type so the byte walkers and the UInt32 codepoint builder share it). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ne guards Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Only the sub-match differs; the negated / keep-restore tail is shared. The backtracker picks the sub-match with a comptime if, so each state still elaborates exactly one arm. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…pers into edfa_match_at/sheng_match_at Every caller passed the DFA's own start states; the LF wrappers passed lf.d's, so they were edfa_match_at/sheng_match_at over lf.d. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
comptime_regex_stages, comptime_pattern_probe, comptime_stages and compile_dashboard each wrote a source file, timed `mojo build` and checked the return code. tools/timed_build.py now does that once; each tool keeps its own flags, cache policy, timeout and error reporting. The pattern probe stays separate: it builds against the precompiled package and ranks many patterns, where the stages ladder builds against the source tree. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
One _add_state call for CHAR/ANY/CHARSET, one MATCH scan for the pinned and unpinned modes (the cut only when unpinned); max_pos was never passed by any caller. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
_sheng_walk_impl and _sheng_full_match_impl were _edfa_walk_impl and _edfa_full_match_impl with a shuffle step (and a state-vector resync after a region skip); cap == 0 keeps the table step. _sheng_step moves to simd_kernels next to the tbl lookups it wraps. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
sweep_ctx superseded them (MULTIPATTERN_PLAN.md: the region-bounded oracle was unsound for lookahead and anchors). __main__ now runs sweep_ctx; _pin keeps a leading global flag group ((?i), (?s)) outside its wrapper, which Python otherwise rejects. sweep_ctx_som stays as the SOM oracle and takes over sweep_som's doc references. Case-table output is byte-identical, and sweep_ctx_som matches the old sweep_som on every case. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The per-start walk moves into an @always_inline _longest_end that both match_at and search_forward inline, so neither gains a call. The unreachable pos > input_len break goes with it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The same bench_set table parse and the same nine row pairs lived in both scripts; the haystack size now comes from the row name's _16k/_64k suffix. A failing `pixi run bench_set` now prints an error and yields no rows (bench_compare_set's handling) instead of a traceback. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
`if "bench" in root` matched the absolute walk root, so a checkout under any path containing "bench" skipped every test (test/ has no bench dir). needs_pkg_rebuild (a != wrapper) and prune_results (one dict comprehension, one caller) are inlined; their unit tests go with them. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- bench_compare: run_mojo_static_benchmarks absorbs _run_mojo_task (its only caller); comments naming the nonexistent bench_static.mojo now name bench.mojo. - unused: statistics (comptime_stages), TA_CENTER and LIGHT_GREEN (gen_pdf). - pixi.toml: the `build` task (mojo precompile -> emberregex.mojoc) was referenced nowhere, duplicated run_test.ensure_package and read like `pixi build`. The lock is unaffected (pixi lock --dry-run --check). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
LFDFA.valid always equalled d.valid (_edfa_finish sets it) and prev_ids only fed one structural test; the look-behind states it appended to `starts` were only renumbered, never read by the finish, so the table is unchanged. That test is deleted: the invariant lives in the builder's look-behind lane, not in the table, and test_both_class_atom_before_anchor / test_differential_both_class_atoms pin the same fix (bd077fe) through the directly built LF table against the Pike VM. build_lf_dfa's minimize hook had no other user. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…able_arr) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
_flatten_nfa's 13 arguments become a _FlatNFA built in place (its constructor also runs _byte_classes); _class_ranges, _word_classes, _eol_ok_bits and _bs_members replace the per-builder copies of the rep_lo/rep_hi, word-class, pending-EOL and bitset->member-list loops. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…alls) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
# Conflicts: # emberregex/engine.mojo # emberregex/onepass.mojo # emberregex/static_dfa.mojo # emberregex/static_lfdfa.mojo # test/test_leftmost_first_dfa.mojo # test/test_word_boundary_dfa.mojo
set_pike_scan and set_pike_som_scan share _pike_scan[som]. The SOM lane's per-thread / per-report starts move out of the recursive closure: every state one closure adds carries the caller's start, so the parallel lists are resized after _set_add_state returns. The plain instantiation keeps the unchanged _set_add_state and _flush_reports (asm-diffed against the base: identical save one cmp/b.ne -> eor/tbnz pair). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Rose's _rose_pos_masks re-derived _litset_pos_masks from the flat meta pools and rose_scan copied litset_scan's chunk loop. The masks now come from litset_masks(Self._rose.lit) at decl level, and both scans run teddy_front_end with their verify step as an inlined closure over an immutable view of the input. The lane probe's teddy/rose functions and constant pools are asm-identical to the previous build. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
bitnfa_scan and bitnfa_stream_chunk repeated the per-byte gather / consume / exception table / shift / seed body; both now call the @always_inline _bn_consume and _bn_advance helpers (block mode still does not call the out-of-line stream function). Lane-probe asm and constant pools are identical to the previous build. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
build_union_subset_nfa and splice_nfa give every pattern a fresh entry state (each fragment start is a state its own build appended), so the dup/-2 detection in _start_id_map and the exact _start_ids scan it selected could only be reached by hand-editing pattern_starts, which is all test_reverse_dfa_shared_entry_uses_exact_start_scan did. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
_rose_walk never read meta, lits or bcls, and _rose_confirm took lits and bcls only to pass them on. Lane-probe asm unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
set_literal's _sort_reports + inline dedup and set_rose's sort_reports + dedup_reports were the same insertion sort (identical predicate) and the same compaction; both lanes now call the Rose pair, moved to set_pike.mojo. The moved bodies are asm-identical; the Teddy lane now calls them out of line instead of inlining its private copy. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
needs_som never read its flags param; has_semantics re-inlined any_ext; sem_table_arr's `b + SEM_MIN_LEN >= n` break could not fire (its only caller sizes n = sem_table_len = SEM_STRIDE * num_patterns). apply_semantics / apply_semantics_spans stay separate: SetMatch has no start field, so one generic needs a trait or per-branch rebinds, which is no shorter than the two 10-line loops. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- _SBN_* / _stream_bn / _can_stream only renamed _BN_* / _bitnfa / _bitnfa.valid; set_stream.mojo and test_set_phase6 read those. - rebind[NFA](self._nfa) was a no-op now that _nfa is a plain NFA. - mdfa_table_str / rdfa_table_str / rose_table_str were pure table_bytes wrappers, rose_flags_arr / ac_cls_arr pure int_arr copies; the decl-level sites (and the bench/tests that used mdfa_table_str) call table_bytes / int_arr directly. Lane-probe asm and static data are identical to the previous build. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
extract_literal_chains expands each state at most once (the seen set), and an expansion pushes at most two successors, so the walk pops at most 2 * num_states + 1 entries; the 2 * num_states + 8 budget could never run out. walk_pops stays (test_set_ac pins it). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
_emit_bits hand-rolled an insertion sort over the accepting ids before its dedup; use sort(ids). Runtime, once per accepting byte on the bitnfa block and stream lanes: revert if bench_set's bitnfa / stream rows regress (the ids lists are small and near-sorted, where insertion sort can win). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
_flush_reports hand-rolled an insertion sort over the ids reported at a position; use sort(ids). Runtime, once per position on the tagged Pike lane (word-boundary sets, residual Pike): revert if bench_set's word-boundary scan row regresses. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
_flush_spans insertion-sorted two parallel lists (ids, starts); emit the spans first and sort that slice of `out` by id with stdlib sort instead (ids are unique per position, so the order is unambiguous). Runtime, once per position on the SOM-carrying Pike (scan_som / scan_spans of word-boundary and reverse-cap-blowup sets): revert if those rows regress. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Both hand-rolled insertion sorts (per-id start/longest order, then the cross-id start/id order) become sort(list, cmp). Both orders are total over their inputs, so stability is moot. Runtime, once per scan_spans call: revert if the scan_spans rows regress. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… _pat Routing Regex.nfa through the _pat ternary put that expression into every lane method's type (via _num_slots) and cost ~20% compile time on bench.mojo. nfa = _build_static_nfa(pattern, flags.value) keeps the default-flags instantiations' expression the same shape as before; _pat stays only at the backtracker call sites, where it measured free. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…UREMENT)" This reverts commit 8805f07. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
apply_flags only ran in the comptime interpreter, so the coverage gate counted its 26 runtime lines as uncovered (98.65% < 99.01% baseline). Three runtime asserts on the existing flags test bring it back to 99.03%. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Benchmark comparisonAdvisory: this check reports but never fails the build. Hosted-runner noise reaches +200% on rows that changed nothing, so treat a flag as a prompt to measure locally, not as a verdict. Runner:
✅ No row regressed more than 20% on both statistics above the 0.02 ms noise floor. Geometric mean of the paired-round medians: -0.1%. All 126 shared rows
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.