ProofGate is a portable evidence contract for agents that generate software. It reduces exhaustive manual review by requiring generated work to demonstrate correctness through a contract, tests, quality gates, metrics when available, adversarial checks, and an explicit verdict.
The portable skill is an instruction package, not an agent runner or deployment system. The repository includes a limited evaluator-side fixture runner; the core package has no runtime dependencies and does not modify host configuration.
The objective is not to make an agent write more code or to replace engineering judgment with a checklist. The objective is to surround generated changes with enough executable evidence and restrictions that unsupported confidence cannot be mistaken for approval.
ProofGate is a stable, maintenance-oriented instruction package. Its supported surface is the portable skill contract, the four operations, the templates, the public evaluation fixtures, and the bounded evaluator runner. New runtime, policy, or host-integration features require evidence from a real workflow; they are not prerequisites for stability.
ProofGate is designed for Windows and Linux. GitHub Actions is configured to run the repository contract suite on both platforms. Development primarily takes place on Windows, while Linux subject workflows have also been validated in isolated Docker environments. macOS is not part of the current scope: its native behavior has not been implemented or validated here.
When a test depends on permissions or behavior specific to another operating
system, ProofGate records that evidence as unavailable (BLOCKED) rather than
presenting the evaluation as complete. Windows remains the project's reference
platform, with Linux as the second supported validation environment.
If a target project has no adequate tests, ProofGate requires the agent to assess the real behavior and add the smallest ecosystem-standard test setup and evidence needed by the contract. Existing green tests are not accepted at face value when they cannot detect plausible defects. Coverage, mutation, static analysis, security, or performance checks are selected when they detect a material risk, not added as decoration.
skills/proofgate/SKILL.md: the host-independent engineering contract.commands/: prompts forplan,build,verify, andauditoperations.templates/: reusable contract and evidence-report formats.evals/: public fixtures and recorded validation evidence.evals/runner.py: evaluator-side workspace preparation, inventory, and gate execution.tests/: contract tests for the package itself.AGENTS.md: repository rules that keep development focused on evidence.
- The reproducible 10-scenario comparison resolved 9/10 tasks without ProofGate and 10/10 with it, reducing critical false success claims from one to zero. See the effectiveness report.
- The first real-repository application found a falsy-item handling defect in
scrapy/queuelib. The public report led to issue #88, and the correction was merged upstream in PR #89. See the PG-R04 evidence. - PG-R09 validated the evaluator runner against a small JavaScript subject with
a prepared defect, demonstrating visible-green/reference-red rejection and a
regression-first final
PASS. - PG-R11 applied the workflow to Gitleaks' unreadable-file handling. It showed
that normal tests can pass while a candidate still fails formatting and its
most important permission-based acceptance test is unavailable on Windows;
ProofGate therefore recorded
FAILwith the portability limitation instead of inferring approval. See the PG-R11 report.
PG-R12 applied ProofGate to bat, a medium-sized Rust command-line project with real CLI, filesystem, and user-output boundaries. Its formatting, build, lint, test, and adversarial CLI checks passed in an isolated Linux environment.
The audit found the tested behavior correct and did not identify a bounded
defect that justified changing the external project. It therefore recorded a
no-change PASS for the audit contract instead of inventing a contribution
task. See the executed evidence and limitations in the
PG-R12 report.
| Operation | Purpose | May edit the target project? |
|---|---|---|
plan |
Define the contract and evidence design | No |
build |
Implement an authorized change and verify it | Yes |
verify |
Check existing work without fixing it | No |
audit |
Find weaknesses and missing evidence | No |
Intensity (lite, full, ultra) and the operational infra profile are
independent selections. See the usage guide.
ProofGate is installed through the host's skill mechanism. For OpenCode, add
this repository's skills/ directory to skills.paths and copy the four
command files from commands/ to a supported command directory. The repository
does not change global configuration automatically.
Example opencode.json entry:
{
"skills": {
"paths": ["/path/to/ProofGate/skills"]
}
}Replace the example path with the cloned repository's skills/ directory.
Other hosts can load the skill file through their supported instruction or skill mechanism and use the same operations in natural language.
The portable skill has no runtime dependency. The evaluator runner and repository contract tests require Python 3.11 or newer and use only its standard library. CI covers Python 3.11 and 3.12 on Windows and Linux:
python -m unittest discover -s tests -vThe test suite validates the package layout, command boundaries, templates, public evaluation fixtures, and reproducibility checks.
The evaluation runner automates fixture mechanics without launching agents:
python evals/runner.py prepare PG-E01 <workspace>
python evals/runner.py inventory <workspace>
python evals/runner.py evaluate PG-E01 <workspace>evaluate returns process status 0 for PASS, 1 for FAIL, and 2 for
BLOCKED. It cannot turn missing manual evidence into a pass.
External repositories may be used as experimental subjects to test ProofGate. They are not part of this package or its dependency graph.
The latest pilot did not add a runtime feature or change the portable skill. The improvement is in the evidence: ProofGate now records a concrete cross-ecosystem case where green tests were insufficient, separates a known format failure from an unavailable permission test, and preserves the reason for the final non-pass verdict.
MIT. See LICENSE.