Skip to content

About

A benchmark of seven PDDL planners that verifies results instead of trusting the service summary line: 6/7 reported success, 2/7 plans that exist. Zero dependencies, 61 tests.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

planning-planner-benchmark

A benchmark of seven PDDL planners on one Depot instance, run through the planning.domains solver service.

The interesting result is not which planner won, but how little the service's own summary line can be trusted. Four of the seven report "Planner found 1 plan(s)" and emit no plan at all; one found three plans and reported none; and the Metric printed for the two that did work is not a plan cost but the last of the timestamps the output formatter invented. Counting the summary line gives 6/7 coverage. Counting plans that actually exist gives 2/7.

The project exists to show how a benchmark harness should verify results rather than trust them.

Zero dependencies. 61 tests.

What the run looks like

$ python examples/compare_logs.py
planner               outcome             steps   seconds  notes
------------------------------------------------------------------------------------
lama-first            solved                 35     5.644  reported Metric 0.034 is the last fabricated timestamp
dual-bfws-ffparser    solved                 68     3.592  reported Metric 0.067 is the last fabricated timestamp
enhsp                 claimed, empty          -     5.286  summary claims 1 plan(s) but the log contains no plan steps
forbiditerative       no plan                 -    16.004  log shows 3 plan(s) found internally that never reached the
metric-ff2            claimed, empty          -     3.614  summary claims 1 plan(s) but the log contains no plan steps
optic                 claimed, empty          -    33.482  summary claims 1 plan(s) but the log contains no plan steps
tfd                   claimed, empty          -     3.407  summary claims 1 plan(s) but the log contains no plan steps

coverage 2/7 by emitted plan, 6/7 by the service's summary line (4 overcounted)

dual-bfws-ffparser is faster (3.592s vs 5.644s) and returns a plan 1.94x longer
(68 vs 35 steps).
On one instance that is an observation, not a benchmark result: n=1, no repeats,
no timing control, shared hosted hardware.

The five decisions worth discussing

1. Coverage counts plans, not claims. The service prints Planner found 1 plan(s) in 5.286secs. for ENHSP, whose log contains zero plan steps and Metric: 0. Any scraper that reads the summary line — which is the obvious thing to write, and the line the service puts there precisely to be read — scores that as a solve. On this instance the two rules differ by four runs out of seven. Verdict.solved is defined as "a plan was emitted" and the report prints both numbers so the gap is visible rather than resolved silently.

2. A plan found but not reported is not counted either. forbiditerative's log says found 3 plans three times and then Planner found 0 plan(s). It is tempting to score that as a solve — it evidently solved the thing. It is not counted, because something between the planner and the output is broken and a comparison's job is to say so, not to repair it. The note is printed.

3. The reported Metric is a rendering artifact, and that is checked, not assumed. The service stamps each plan step with a timestamp incrementing by 0.001 and then reports the final one as both Metric and Makespan. For the 68-step plan that is 0.067; for the 35-step plan, 0.034. Comparing those numbers across planners compares plan lengths in disguised units and would be meaningless if any action had a cost other than 1. metric_is_a_timestamp_artifact tests the value numerically, so a planner whose Metric genuinely carries cost is left alone — asserting the artifact without checking would be the same error inverted.

4. Failures sort by name, never by time. OPTIC failed in 33 seconds and TFD failed in 3.4. Ordering failures by speed puts TFD near the top of the table and invites reading it as a ranking. A planner that failed quickly did not win anything.

5. n=1 travels with the result. Two runs produced plans, on one instance, on shared hosted hardware, with no repeats and no timing control. The quality/speed sentence says so in the output itself rather than leaving it to whoever quotes the table. The 1.94x plan-length ratio is a real observation about these two runs and evidence of nothing about either planner in general.

What the comparison is actually good for

The one result that survives the caveats is the shape of the trade-off: dual-BFWS returned a plan 1.9x longer in 64% of the time. That is the familiar satisficing-planner trade — the first plan a greedy best-first search finds is not the cheap one — and it is visible here because plan length was recomputed from the emitted steps rather than read from a field named Metric.

Design

src/plannerbench/
  parse.py     log -> RunLog. Extracts, never interprets.
  verdict.py   RunLog -> Verdict. The coverage rule and the artifact check.
  compare.py   Verdicts -> the table, the two coverage numbers, the caveats.
  cli.py       argument parsing and the --strict exit code.

The split between parse and verdict is the one that matters: the parser records what the log says, including the claims that turn out to be false, and every judgement lives in one small module that can be read in a minute and argued with. A parser that dropped reported_plan_count because it is unreliable would also have destroyed the finding.

Usage

python3.12 -m venv .venv
source .venv/bin/activate          # Windows: .venv\Scripts\activate
pip install -e ".[dev]"
pytest -q
python examples/compare_logs.py
python -m plannerbench                       # every log under logs/
python -m plannerbench logs/lama-first.txt
python -m plannerbench --strict              # exit 1 if any run claims a plan it did not emit

--strict exists so this is usable as a check on a pipeline that submits jobs to the service: a run that claims success and returns nothing should fail a build, not appear in a results table.

As a library:

from plannerbench import coverage, judge, parse_log_file

verdicts = [judge(parse_log_file(p)) for p in Path("logs").glob("*.txt")]
print(coverage(verdicts))
for v in verdicts:
    print(v.run.planner, v.outcome.value, v.notes)

Dataset

None to download. The inputs are in this repository:

File What it is
domain.pddl The Depot domain — trucks, hoists, crates, pallets
instance-14.pddl The instance every log in logs/ was run on
instance-18.pddl A larger instance, kept for re-running
logs/*.txt Raw output from each solver, unedited

To reproduce or extend: paste domain.pddl and an instance into https://editor.planning.domains, choose a solver package, and save the output as logs/<solver-name>.txt. The planner name is taken from the file stem, so name the file after the solver package.

Every log in logs/ is unedited, including the four that claim a plan they did not produce. Cleaning those up would delete the finding.

Not included

  • No timing methodology. Wall-clock seconds come from the service's own summary line, measured on shared hardware under unknown load. They are reported because they are what exists, and they carry the n=1 caveat.
  • No multi-instance coverage table. Seven planners on one instance is a case study. A real coverage benchmark needs the IPC instance set, a local build of each planner, and a fixed time and memory limit per run.
  • No plan validation. Whether the 68-step and 35-step plans actually achieve the goal is a separate question, answered by the validator in the sibling repository 08.28.planning-pddl-rovers.
  • No repeats or confidence intervals. One run each. With n=1 an interval would be decoration on a number that has no distribution behind it.

License

MIT — see LICENSE.

About

A benchmark of seven PDDL planners that verifies results instead of trusting the service summary line: 6/7 reported success, 2/7 plans that exist. Zero dependencies, 61 tests.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages