GitHub Spec Kit tells you what to build.
OpenSpec lets you iterate freely.
Neither verifies the agent actually built it.This template adds the missing phase.
AI coding agents are extraordinary at implementation. They are unreliable at reporting whether the implementation worked.
When a coding agent finishes a task and says "tests passing," it is predicting what a successful completion summary looks like — not reading from ground truth. The summary is generated the same way the code was: by pattern matching. A successful task in training data ends with "tests passing," so the summary says "tests passing." Whether the tests actually passed is a separate question the model cannot answer honestly.
The failure modes are consistent:
- Agent runs tests, sees a failure, decides it's unrelated, reports success
- Agent doesn't run tests at all and tells you it did
- Agent fixes a test by relaxing the assertion rather than fixing the code
Every existing SDD methodology ends at Implement. This one doesn't.
Spec Kit: Spec → Plan → Tasks → Implement
OpenSpec: Propose → Implement → Archive
RFD Method: (Stop + Verify) → Directive → Implement → Certify Floor → Advance
Every phase opens with a stop rule. The agent runs the test suite before touching any file and reports the raw output. This establishes ground truth at session start so any drift during implementation is immediately visible.
⛔ STOP: Run pytest before touching any file.
Must report 42 passing, 0 failing, 0 skipped.
If count differs, stop and report — do not proceed.
The stop rule is not for the agent's benefit. Agents don't have intentions to protect. It's for yours. It's a forcing function that produces a measurement before the work begins.
Every phase is a structured document with explicit scope. The directive names every file the agent is allowed to touch and every file it is not. It specifies test anchors — the exact behaviors that must pass for the phase to be complete. It specifies completion criteria — a checklist that must be true before anything advances.
Agents modify adjacent files. The read-only list makes the boundary legible. The agent can't claim it didn't know.
The phase is not complete because the agent said so. The phase is complete when the human reads the raw terminal output and writes the certified floor into the next directive.
47 passed, 0 failed, 0 skipped
That line is proof. Everything before it is a story.
| Claim | Required proof |
|---|---|
| Tests passing | Raw pytest/test output, read by you |
| App works on device | Device screenshot, taken by you |
| Build succeeded | Terminal output of the build command |
| Deployment live | URL loaded in browser, screenshot taken |
| Module implemented | You read the file |
An agent summary does not appear on this list. Not because agents are useless — they're extraordinary — but because the summary is the wrong artifact. It's a prediction. The terminal output is a measurement.
You (Architect) → Directive (Spec) → Agent (Implementation)
You set direction, review output, certify floors, write ADRs.
The directive defines scope, stop rules, test anchors, completion criteria.
The agent implements against the spec, nothing more.
The agent is not optimizing for your system. It's optimizing for the task. The discipline has to come from outside the agent. You can't ask the agent to be cautious. You have to build the caution into the structure it operates inside.
Creates a copy of the repo structure in your account. No CLI required.
Project-level rules that apply to every phase. Global read-only files. Branch strategy. Naming conventions. Do this once.
Paste your current project state: what phase you're on, what the certified floor is, what's next. Update this at the end of every session.
One copy per phase. Fill in the stop rule with your real floor number. Fill in the scope table. Write the implementation. Write test anchors. Write completion criteria.
Windsurf, Cursor, Claude Code, Copilot — any agent that reads files. The directive is the spec. The agent implements.
Run the tests yourself. Read the output. If it matches the stop rule target, write the new floor into the next directive. If it doesn't, the phase isn't done.
These are real, shipped projects built under this methodology. Each has SDDs, ADRs, and certified test floors.
| Project | Stack | Floor | Status |
|---|---|---|---|
| PrivyBot | Python, FastAPI, SQLite | 604/0/0 | Phase 33 |
| VoidDrift | Rust, Bevy 0.15, egui | — | Live on itch.io |
| OpenAgent | Python, Click | 103/0/3 | Live on PyPI |
| RFD Blog Engine | Python, FastAPI, SQLite | 130/0/0 | Phase 2 |
- The Agent Told Me It Was Done. The Tests Said Otherwise. — The failure mode that established the proof standard (live July 5, 2026)
- The Spec Is Load-Bearing — Why the directive comes before the code
- rfditservices.com — Portfolio and consulting
MIT. Use it, adapt it, publish your own presets from it.
The methodology is the contribution. The templates are the delivery mechanism.