arc-bench is a benchmark for requirement-to-application generation. It
evaluates whether a generation system can transform multi-modal web application
requirements into a runnable implementation whose behavior is validated by
end-to-end Playwright tests.
The benchmark is organized as a set of web application tasks. Each task pairs a requirement package with an executable test suite, so different generators can be compared against the same inputs and behavioral checks. This repository contains:
arc-bench/webapp/<app>/requirements/: structured requirements for the benchmark web apps;arc-bench/webapp/<app>/tests/: Playwright tests for those apps;scripts/run-playwright.jsandplaywright.config.ts: the benchmark test runner;Dockerfile: an optional containerized benchmark execution environment.
The benchmark itself is generator-agnostic: any method can consume the requirements, produce a web application, start it locally, and run the provided tests against the application URL.
Requirement counts are the number of atomic requirement nodes in
requirements.yaml. Test counts are the number of Playwright test(...) cases.
| App | # Requirements | # Test cases | # Domain |
|---|---|---|---|
keep |
32 | 32 | Google Keep, https://keep.google.com/ |
bookstack |
34 | 34 | BookStack, https://demo.bookstackapp.com/ |
stackoverflow |
66 | 66 | Stack Overflow, https://stackoverflow.com/ |
prestashop |
86 | 86 | PrestaShop, https://demo.prestashop.com/ |
12306 |
117 | 117 | China Railway 12306, https://www.12306.cn/en |
ctrip |
125 | 125 | Ctrip, https://www.ctrip.com/ |
The benchmark usage is independent of any particular generation method:
arc-bench/webapp/<app>/requirements/
-> generate a runnable web application with a chosen method
-> start the generated application
-> run arc-bench/webapp/<app>/tests/ against the application URL
Install the local test runner when running tests directly on the host:
npm install
npm run test:installRun one benchmark application's tests against a running application:
npm run test -- --app bookstack --target-url http://127.0.0.1:3301The Docker image in this repository provides a benchmark execution environment: Node.js, Playwright browsers, the test runner, and benchmark files. For non-ARC generators, you can use the image as a clean test environment by mounting a generated application and running the benchmark tests against it.
This section is an application example of the benchmark using ARC (Agentic Requirement Compiler) as the generation method. This repository is not a standalone implementation of the ARC compiler; the compiler and its application templates are provided through Git submodules.
Install the following for the ARC baseline example:
- Docker Desktop on Windows/macOS or Docker Engine on Linux;
- access to an OpenAI-compatible model API.
Clone the benchmark repository and initialize the ARC compiler submodule and its nested submodules:
git submodule sync --recursive
git submodule update --init --remote --recursiveThe environment file belongs to the ARC compiler submodule. Read the configuration instructions in:
agentic-requirement-compiler/README.md
Create the compiler environment file from the template.
Linux/macOS:
cp agentic-requirement-compiler/.env_example \
agentic-requirement-compiler/.envWindows PowerShell:
Copy-Item `
agentic-requirement-compiler\.env_example `
agentic-requirement-compiler\.envEdit agentic-requirement-compiler/.env according to the ARC compiler README.
At minimum, configure the model API credentials and model name. The file is
passed to Docker at runtime and is excluded from the Docker image.
For one application, the complete flow is:
arc-bench/webapp/<app>/
|-- requirements/ input requirements and reference assets
`-- tests/ Playwright tests for the generated app
ARC compiles `arc-bench/webapp/<app>/requirements/`
-> generated backend starts on port 3301
-> /api/health becomes available
-> Playwright tests run from `arc-bench/webapp/<app>/tests/`
-> application, logs, raw results, and HTML report are exported
The recommended ARC baseline command performs all steps in one isolated
container. The examples below use bookstack; replace it with another benchmark
app name as needed.
If you add or modify arc-bench/, apps.config.json, scripts/, or
docker/entrypoint.sh, rebuild the image before running the container again.
The image copies those files at build time.
Linux:
mkdir -p docker-output
docker run --rm \
--env-file agentic-requirement-compiler/.env \
--mount "type=bind,source=$PWD/docker-output,target=/export" \
arc-reproduction:latest bookstackmacOS:
mkdir -p docker-output
docker run --rm \
--env-file agentic-requirement-compiler/.env \
--mount "type=bind,source=$PWD/docker-output,target=/export" \
arc-reproduction:latest bookstackWindows PowerShell:
New-Item -ItemType Directory -Force docker-output | Out-Null
docker run --rm `
--env-file agentic-requirement-compiler\.env `
--mount "type=bind,source=$((Get-Location).Path)\docker-output,target=/export" `
arc-reproduction:latest bookstackReplace bookstack with one of:
keep
bookstack
stackoverflow
prestashop
12306
ctrip
The container entrypoint performs:
arc compile
-> PORT=3301 npm run start
-> wait for http://127.0.0.1:3301/api/health
-> npm run test -- --app <app-name>
The container exits with:
0when compilation, startup, and tests succeed;- a non-zero status when compilation fails, the health check fails, or tests fail.
The one-container command above is recommended for complete reproduction. The following commands are useful when debugging or running one stage separately.
Linux/macOS:
docker build --progress=plain -t arc-reproduction:latest .Windows PowerShell:
docker build --progress=plain -t arc-reproduction:latest .Parameter meanings:
docker build: builds an image from theDockerfile;--progress=plain: prints complete build logs;-t arc-reproduction:latest: assigns the image name and tag;.: uses the current repository as the Docker build context.
This command runs ARC compilation without starting the generated application or running Playwright.
Linux/macOS:
docker run --rm \
--env-file agentic-requirement-compiler/.env \
--mount "type=bind,source=$PWD/docker-output,target=/export" \
--entrypoint arc \
arc-reproduction:latest \
compile /opt/arc/arc-bench/webapp/bookstack/requirements \
-o /export/bookstack/application \
--type web \
--cleanWindows PowerShell:
docker run --rm `
--env-file agentic-requirement-compiler\.env `
--mount "type=bind,source=$((Get-Location).Path)\docker-output,target=/export" `
--entrypoint arc `
arc-reproduction:latest `
compile /opt/arc/arc-bench/webapp/bookstack/requirements `
-o /export/bookstack/application `
--type web `
--cleanParameter meanings:
--entrypoint arc: bypasses the default Docker entrypoint and calls the installed ARC CLI directly;compile: compiles a requirement directory into an application;/opt/arc/arc-bench/webapp/bookstack/requirements: the requirement directory inside the image;-o /export/bookstack/application: the generated application output directory;--type web: selects web application generation;--clean: removes an existing output directory before compilation;--mount ...:/export: persists the container output underdocker-output/bookstack/applicationon the host.
The default entrypoint already starts and tests a newly generated application. To test an existing generated application, mount it into a fresh container, start its backend, wait for the health endpoint, and run Playwright.
Linux/macOS:
docker run --rm \
--mount "type=bind,source=$PWD/docker-output/bookstack/application,target=/workspaces/bookstack" \
--entrypoint /bin/bash \
arc-reproduction:latest \
-lc 'set -e
(cd /workspaces/bookstack/backend && PORT=3301 npm run start >/tmp/arc-app.log 2>&1) &
server_pid=$!
trap "kill $server_pid 2>/dev/null || true" EXIT
until curl --fail --silent http://127.0.0.1:3301/api/health >/dev/null; do sleep 1; done
cd /opt/arc
TARGET_URL=http://127.0.0.1:3301 npm run test -- --app bookstack'Windows PowerShell:
docker run --rm `
--mount "type=bind,source=$((Get-Location).Path)\docker-output\bookstack\application,target=/workspaces/bookstack" `
--entrypoint /bin/bash `
arc-reproduction:latest `
-lc 'set -e; (cd /workspaces/bookstack/backend && PORT=3301 npm run start >/tmp/arc-app.log 2>&1) & server_pid=$!; trap "kill $server_pid 2>/dev/null || true" EXIT; until curl --fail --silent http://127.0.0.1:3301/api/health >/dev/null; do sleep 1; done; cd /opt/arc; TARGET_URL=http://127.0.0.1:3301 npm run test -- --app bookstack'This form assumes the generated application already contains its dependencies and frontend build output. Otherwise, use the complete one-container command.
Linux/macOS:
for app in keep bookstack stackoverflow prestashop 12306 ctrip; do
docker run --rm \
--env-file agentic-requirement-compiler/.env \
--mount "type=bind,source=$PWD/docker-output,target=/export" \
arc-reproduction:latest "$app"
doneWindows PowerShell:
$apps = @("keep", "bookstack", "stackoverflow", "prestashop", "12306", "ctrip")
New-Item -ItemType Directory -Force docker-output | Out-Null
foreach ($app in $apps) {
docker run --rm `
--env-file agentic-requirement-compiler\.env `
--mount "type=bind,source=$((Get-Location).Path)\docker-output,target=/export" `
arc-reproduction:latest $app
}Each application runs in its own container and does not reuse another application's backend process or workspace.
Results are written to:
docker-output/<app>/
|-- application/ generated ARC application
|-- logs/
| |-- compile.log ARC compilation log
| |-- app.log application startup log
| `-- test.log Playwright log
|-- test-results/ screenshots, videos, traces, and raw results
|-- playwright-report/ Playwright HTML report
`-- summary.txt compilation, runtime, and test statuses
Open the HTML report at:
docker-output/bookstack/playwright-report/index.html