English · 简体中文 · Evidence schema · v1 JSON
Version status: current software is the
v1.1.2maintenance patch. The research protocol and evidence baseline remain the immutablev1.0.0public release.v1.1.2keeps that boundary and fixes checkpoint-label containment and equal row geometry in the Studio; it does not add a new regression claim. See the maintenance policy and portfolio evidence.
Browser Agent Regression is a local-first harness for answering one practical question: did a browser-agent change improve reliability, or did it quietly break a workflow that used to pass?
The v1 core provides three resettable browser tasks, four meaning-preserving UI variants, checkpoint-level independent scoring, repeated baseline/candidate comparison, first-failure localization, and verifiable JSON evidence. The default demo uses built-in deterministic drivers; model-backed adapters can apply the same evidence protocol to real-agent experiments.
Python 3.11–3.13 is supported.
git clone --branch v1.1.2 --depth 1 https://github.com/AlbertXXuu/BrowserAgentRegression.git
cd BrowserAgentRegression
python -m venv .venv
# Windows: .venv\Scripts\activate
# macOS/Linux: source .venv/bin/activate
python -m pip install -e ".[dev]"
python -m playwright install chromium
browser-agent-regression doctor
browser-agent-regression demo
browser-agent-regression studioThe demo runs 12 controlled attempts and writes runs/demo-report.json. The expected result is clean
parity plus one deliberately induced popup regression for each task, localized to the first failed
checkpoint. The report records a synthetic harness-calibration result with its drivers and protocol.
studio opens a loopback visual interface that follows the AlvenX product design language. It starts
from the hash-checked committed evidence and can run the same deterministic calibration in memory
while leaving the committed report unchanged. See Studio usage.
Verify either a new report or the committed v1 evidence:
browser-agent-regression verify --report runs/demo-report.json
browser-agent-regression verify --report docs/evidence/v1.0.0-calibration.json- CLI commands:
demo,oracle,calibrate,verify,doctor,serve,studio, and optionaldeepseek. - Task IDs, variant IDs, and ordered checkpoint contracts documented in packaged JSON manifests.
- Evidence schema
1.0, protocolbrowser-agent-regression-controlled-ui-v1, fixture hashes, environment metadata, per-attempt outcomes, summaries, regressions, and first failures. - Exit codes:
0for a passing command,1for a completed experiment that misses its acceptance condition, and2for invalid evidence or a runtime/setup failure.
The validator recomputes summaries and regressions from attempts, checks every checkpoint contract,
verifies fixture SHA-256 values against the installed package, and rejects credential-shaped fields.
Historical schema 0.2 reports remain readable and verifiable.
Check the reference driver across all tasks and variants:
browser-agent-regression oracle --runs 30 --output runs/oracle.jsonCompare the reference driver with a deliberately popup-blind candidate:
browser-agent-regression calibrate --runs 10 --output runs/calibration.jsonRun one task or variant while diagnosing a fixture:
browser-agent-regression oracle \
--task catalog.find-and-save.v1 \
--variant delayed-render \
--runs 3Read the calibration report in English or 简体中文, with results, first-failure checkpoints, and the separate real-agent feasibility boundary.
The committed v1 calibration repeats each baseline/candidate/task/variant cell three times. The
reference and candidate match on clean pages; the candidate falls from 100% to 0% under the popup
overlay for all three tasks, with 100% agreement on each first failed checkpoint. The companion v1
oracle report covers all four variants. Both reports pass the same public verify command used in CI.
Use the failure taxonomy to separate that first unmet outcome
checkpoint from a causal claim about observation, action, recovery, safety or the harness.
The research landscape compares the frozen v1 boundary with primary
sources for adjacent benchmarks, diff tools, agent evaluation, diagnosis and WebMCP testing.
Earlier Phase 0 evidence is retained as historical provenance, including a one-run Browser Use + DeepSeek feasibility check. That paid adapter showed integration feasibility but not repeated model reliability; the independent DOM scorer remains authoritative.
The Browser Use + DeepSeek adapter applies the same tasks and independent DOM scoring to a model-backed run. It can make paid API requests:
python -m pip install -e ".[agent]"
browser-agent-regression deepseek \
--task preferences.notifications.v1 \
--runs 1 \
--headedRead the setup and safe-key guide first. API credentials are read from the environment or hidden input and are never written into evidence.
v1 stabilizes the harness, CLI, package, and evidence contract for external use. Current work focuses on setup, evidence readability, and integrations for named agents and models. Use the controlled workflow for repeatable regression localization, and combine it with WebArena, BrowserGym, or team evals when broader task coverage is required.
python -m ruff check .
python -m pytest
python scripts/check_repository.py
python -m buildApache-2.0 licensed. Browser Agent Regression is an AlvenX open-source project.
