CodeGraph ships a reproducible comparison harness in codegraph-bench. It is
separate from correctness: a faster result never relaxes schema, golden, or graph
parity.
For each pinned corpus, the harness compares the release Rust CLI with a built copy of the pinned upstream TypeScript CLI on byte-identical source trees. It records:
- cold and warm full-index wall time;
- one-file incremental synchronization wall time;
- cold and warm query latency (median, MAD, p50, and p99);
- peak RSS samples when the host exposes them;
- resulting database size; and
- a Rust-only in-process parse metric, clearly separated because upstream has no equivalent entry point.
Before timing a corpus, the pipeline indexes both arms and requires normalized
SQLite .schema equality. A schema mismatch aborts the run rather than comparing
non-equivalent work.
The source registry in crates/codegraph-bench/src/corpus.rs is authoritative.
Each row pins a repository URL, full commit SHA, benchmark subdirectory, expected
non-blank LOC, and source-file count. --fetch-corpora creates immutable
checkouts under ignored bench/corpora/ and refuses an existing checkout at the
wrong commit.
Inspect the registry and local state:
cargo run --locked -p codegraph-bench --bin bench -- --list-corporaFetch the exact commits (networked, explicit operation):
cargo run --locked -p codegraph-bench --bin bench -- --fetch-corpora-
Build the shipped Rust binary with the same locked release profile used by release workflows:
cargo build --locked --release -p codegraph-rs
-
Materialize the pinned upstream checkout at
reference/colbyand buildreference/colby/dist/bin/codegraph.jsaccording to that checkout'sRUN.md. Record its exact commit in the result artifact. -
Fetch the corpus pins and verify
--list-corporareports the expected commit, file count, and LOC for every selected corpus. -
Use an otherwise idle machine. Record CPU, memory, OS, kernel, Node, Rust, Rust implementation commit, and upstream implementation commit.
The full pipeline is Linux-oriented because peak-RSS collection and cold-cache handling use Linux facilities when available. If the kernel does not allow page cache drops, the harness records best-effort userspace eviction; do not label that series equivalent to a privileged cold-cache run without the emitted caveat.
Use at least two runs because run one is discarded. A normal evidence run should use enough samples for stable p99 interpretation; twelve is the CLI default.
cargo run --locked --release -p codegraph-bench --bin bench -- \
--run \
--runs 12 \
--corpora all \
--out target/benchmark-results.json \
--report-md docs/benchmark-results.mdTo isolate one corpus during harness development:
cargo run --locked --release -p codegraph-bench --bin bench -- \
--run --runs 12 --corpora fd-small \
--out target/benchmark-fd.jsonTo re-render Markdown without rerunning measurements:
cargo run --locked -p codegraph-bench --bin bench -- \
--render-md target/benchmark-results.json docs/benchmark-results.mdA publishable result must include:
- raw JSON and generated Markdown from the same run;
- clean Git status or an explicit diff description;
- both implementation commit SHAs;
- exact corpus pins and observed counts;
- environment fields emitted by the harness;
- run count, discarded-first policy, and cache mode;
- schema equality for every corpus;
- every raw sample, median, MAD, p50, and p99;
- failures, timeouts, unsupported RSS, or cache-eviction caveats.
Do not copy a latency number into README, AGENTS, or release notes without linking to such an artifact. Results are snapshots, not timeless product guarantees.