Skip to content

issue_82: Invoke the Convergence Check from the participating skills - #85

Merged
dieterbaier merged 3 commits into
mainfrom
issue_82
Aug 31, 2026
Merged

issue_82: Invoke the Convergence Check from the participating skills#85
dieterbaier merged 3 commits into
mainfrom
issue_82

Conversation

@dieterbaier

@dieterbaier dieterbaier commented Aug 31, 2026

Copy link
Copy Markdown
Member

Summary

Wires the canonical Convergence Check from #84 into the three skills that own the moments it applies to, and specifies the wiring itself. A gate nobody invokes is documentation, not a gate.

Closes #82

Where it lands, and why there

implement-issue-workflow — three touch points, because a result describes one state of a branch, not the branch:

  • Commit, Push, And PR — an optional run, recorded as explicitly provisional.
  • Address PR Comments — a recorded result is void once new commits land; replace it, do not amend it.
  • PR Integration — the authoritative run, against the commit that would be integrated. A result that does not cover the current head joins the existing warning-signs list, alongside failed checks and unclear branch ownership. That list already collects the reasons not to integrate automatically, so the gate belongs inside it rather than beside it as a second mechanism.

pr-review — a step to check the reported result against the canonical definition, plus a Required Reading entry so the review compares against the file rather than a remembered version. A result claimed without its evidence, a finding without a disposition, or a blocker without its kind is itself a review finding; and when a pull request reports no result, the reviewer says so rather than producing one — a reviewer generating the result they are meant to check is not a check.

architecture-impact — a final workflow step and a Review Checklist line. Impact analysis establishes what a change touches; the gate asks whether the artifacts still agree afterwards.

The wiring is specified, not asserted

These skills are the executable behaviour contracts for agents, so this change does alter observable behaviour: agents must now run and record the gate, and must not auto-integrate on certain results. features/skill-wiring.feature and test/skill-wiring.test.mjs specify what is verifiable about it — four scenarios:

  • every skill-to-skill reference resolves (53 across the tree);
  • the three callers reach the gate;
  • no caller defines any of the four result states;
  • no caller carries the seven-question structure.

Both uniqueness guards read what to look for from the canonical file rather than hard-coding it, so renaming a state or a question keeps the guard pointed at the right thing instead of quietly guarding nothing.

Each guard was mutation-probed rather than assumed:

Mutation Result
reference bent to convergence-checkX not ok 1 — Every skill-to-skill reference resolves
| **Blocked** | … pasted into pr-review not ok 3 — skills/pr-review/SKILL.md defines "Blocked"
four question headings pasted without their section heading not ok 4 — skills/pr-review/SKILL.md carries 4 question headings

One rule moved into the canonical skill

Applying the gate to this pull request exposed a gap in it. I first reported the absent execution layer for prose contracts as an unavailable blocker — but it would appear identically on every change to any prose skill, so blocking on it would stop all such work until #65 lands while saying nothing about the change under review.

skills/convergence-check/SKILL.md now separates the two: a blocker is evidence this change owes and cannot produce; a limitation of the medium or the tooling is a residual risk, which question 6 already requires reporting together with its follow-up. The test is one question — would this same blocker appear on every change of this kind? A gate that can never be passed is not a gate.

No rule is duplicated

Verified by the new tests, not by inspection: the four result-state definitions and the seven-question structure appear in exactly one file. This is the check's own rule applied to itself — a second copy of a rule inside a gate is the drift the gate exists to detect.

Convergence Check — authoritative run against 34f9c0d

  • Result: converged.
  • Intent and scope — the three edits map to Slice #77.3: Invoke the Convergence Check from the participating skills #82's acceptance criteria; the adapter criterion was satisfied on arrival by issue_81: Define the canonical Convergence Check #84.
  • Behaviourapplicable and specified. Observable agent behaviour changes, and the verifiable part carries a Gherkin specification bridged to four automated tests, each mutation-probed.
  • Architecture impact — no boundary, component, interface or deployment element changes; skills and their tests only.
  • Decisions — one decision recorded in the canonical skill (standing limitation versus change-specific blocker) rather than left in this discussion.
  • Traceabilityfeatures/skill-wiring.feature names its bridge; the three references resolve and are tested.
  • Implementation and verification./build.sh test 39 JS tests, 157 Ruby assertions, 0 failures; ./build.sh check-adapters current. No new skill, so adapter output is unchanged.
  • Documentation and delivery state — this description now matches the diff at head 34f9c0d. An earlier revision did not: it still claimed Behaviour — not applicable, no executable path, and 35 tests, after the change had grown a specification and four more. Question 7, found on review — the same finding this gate caught on issue_80: Present the toolkit through capabilities #83.
  • Residual risk, reported per question 6 — whether an agent obeys a prose contract is not verifiable in this repository. Standing limitation of the medium; follow-up [EPIC] Establish Agent Skills conformance and evaluation #65. Not a blocker, not a waiver.

A gate nobody invokes is documentation. Three skills own the moments it
applies to, and each now names the moment and defers everything else.

- implement-issue-workflow runs it before declaring a pull request ready
  and records the result in the PR body. A result other than converged or
  converged with recorded waivers joins the existing warning signs that
  stop automatic integration — the list already collects reasons not to
  merge, so the gate belongs in it rather than beside it.
- pr-review checks the reported result against the canonical definition.
  A result claimed without its evidence, a finding without a disposition,
  or a blocker without its kind is a review finding; a missing result is
  reported rather than supplied from the review, because a reviewer
  producing the result they are meant to check is not a check.
- architecture-impact runs it before work is treated as complete. Impact
  analysis establishes what a change touches; the gate asks whether the
  artifacts still agree afterwards.

None of the three restates a question, a result state, or an evidence
rule. A second copy of a rule inside a gate is the drift the gate exists
to detect.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01USXp58FoRppFK6FUht8K7u

@dieterbaier dieterbaier left a comment

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Vielen Dank für die saubere, bewusst schlank gehaltene Verdrahtung. Die Referenzen lösen korrekt auf, die kanonischen Regeln werden nicht dupliziert und die CI ist grün. Zwei zusammenhängende Punkte verhindern aus meiner Sicht aber noch, dass das Gate seine beabsichtigte Wirkung zuverlässig entfaltet:

  1. Der Convergence Check wird im Implementierungsworkflow zu früh ausgeführt.

    Der neue Schritt liegt vor „declaring the pull request ready“. Der kanonische Skill verlangt den maßgeblichen Lauf dagegen am Ende, bevor der PR als mergeable gilt, und stellt ausdrücklich klar, dass ein früher Lauf noch kein abschließendes Ergebnis ist. Nach dem Ready-Zeitpunkt können Review-Kommentare, weitere Commits und zusätzliche Verifikationen den geprüften Stand verändern. Vor der Integration wird das vorhandene Ergebnis lediglich bewertet; ein erneuter Lauf oder eine Prüfung seiner Aktualität ist nicht vorgesehen.

    Bitte den Check nach der Bearbeitung der Review-Kommentare beziehungsweise unmittelbar vor der Integrationsentscheidung erneut ausführen lassen und das dokumentierte Ergebnis aktualisieren. Erst dann ist auch das Akzeptanzkriterium aus #82 („before a feature PR is declared mergeable“) zuverlässig erfüllt.

  2. Die Selbsteinschätzung „Behaviour — not applicable“ ist nicht schlüssig.

    Der PR verändert beobachtbares Agentenverhalten: Agenten müssen künftig den Check ausführen und dokumentieren und dürfen PRs mit bestimmten Ergebnissen nicht automatisch integrieren. Die geänderten Skills sind hier gerade die ausführbaren Verhaltensverträge. Dass kein klassischer Codepfad geändert wird, macht die Verhaltensfrage daher nicht „not applicable“.

    Bitte den Convergence Check für dieses Verhalten tatsächlich durchführen und die passende Evidenz beziehungsweise – falls eine BDD-Spezifikation für diese Änderung unverhältnismäßig ist – eine begründete, von einem Menschen akzeptierte Waiver dokumentieren.

Abgesehen davon habe ich keine weiteren Findings: Die relativen Referenzen passen, die Schritte im PR-Review wurden vollständig neu nummeriert, und ich sehe keine unerwünschte Kopie der sieben Fragen oder der Ergebniszustände.

Review on #85 found two connected problems.

The check ran too early. The step sat before "declaring the pull request
ready", while the canonical skill puts the authoritative run before
mergeability and says plainly that an early pass is not a result. Review
comments and later commits change the state it inspected, and PR
Integration only read the recorded result without asking whether it was
still current.

The run at PR-open time is now explicitly provisional. Address PR
Comments voids a recorded result when new commits land, and PR
Integration runs the authoritative check against the commit that would be
integrated. A result that does not cover the current head joins the
warning signs, because a stale result is not a result.

The second finding was mine to own: this change alters observable agent
behaviour — agents must run and record the gate, and must not integrate
on certain results — so reporting Behaviour as "not applicable" was
wrong. The skills are the behaviour contracts here.

What can be specified now is, and it is bridged: skill-to-skill
references must resolve, the three callers must reach the gate, and its
result-state definitions and question structure must live in exactly one
file. A mutation probe confirms the guard fails when a reference is
broken.

Writing that guard sharpened a distinction the prose had left implicit.
Naming a result state elsewhere is legitimate — implement-issue-workflow
says which results stop integration, which is its own policy in the
gate's vocabulary. Defining one elsewhere is the copy that drifts, so the
test targets the definition, not the word.

What remains unverifiable here is whether an agent obeys prose at all.
That is #65, and it is reported as an unavailable blocker rather than
waived.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01USXp58FoRppFK6FUht8K7u
@dieterbaier

Copy link
Copy Markdown
Member Author

Beide Punkte nachgearbeitet in 7067063. Der zweite war der wichtigere — und meiner.

1. Der maßgebliche Lauf sitzt jetzt vor der Integration

Du hast die Lücke genau beschrieben: zwischen „ready" und Integration landen Review-Kommentare, weitere Commits und zusätzliche Verifikation, und PR Integration hat das aufgezeichnete Ergebnis nur gelesen, ohne nach seiner Aktualität zu fragen.

  • Commit, Push, And PR — der Lauf dort ist jetzt ausdrücklich provisorisch markiert.
  • Address PR Comments — ein aufgezeichnetes Ergebnis ist ungültig, sobald neue Commits landen. Ersetzen, nicht ergänzen: ein Ergebnis beschreibt einen Stand des Branches, nicht den Branch.
  • PR Integration — der maßgebliche Lauf findet hier statt, gegen den Commit, der integriert würde. Ein Ergebnis, das den aktuellen Head nicht abdeckt, ist jetzt selbst ein Warnzeichen. Ein veraltetes Ergebnis ist kein Ergebnis.

2. „Behaviour — not applicable" war falsch

Du hast recht, und mein eigener Skill hätte mich korrigieren müssen: Die geänderten Skills sind die ausführbaren Verhaltensverträge. Agenten müssen den Check künftig laufen lassen, dokumentieren und dürfen bei bestimmten Ergebnissen nicht automatisch integrieren. Das ist geändertes beobachtbares Verhalten.

Statt einer Waiver habe ich den spezifizierbaren Teil spezifiziert und gebrücktfeatures/skill-wiring.feature mit test/skill-wiring.test.mjs, vier Szenarien:

  • jede Skill-zu-Skill-Referenz löst auf (53 Referenzen über alle Skills);
  • die drei Aufrufer erreichen das Gate;
  • die Definitionen der Ergebniszustände liegen in genau einer Datei;
  • die Fragenstruktur ebenso.

Mutationsprobe statt Behauptung: Ich habe eine Referenz in pr-review auf convergence-checkX verbogen — Test 1 schlägt fehl, danach aus git wiederhergestellt. Der Wächter hat einen Fehlerpfad, was nach dem Skill selbst die Bedingung dafür ist, dass er zählt.

Das Schreiben des Tests hat eine Unterscheidung geschärft, die in der Prosa implizit blieb. Der erste Entwurf schlug fehl, weil implement-issue-workflow zwei Ergebniszustände beim Namen nennt. Das ist aber legitim: Der Skill sagt, welche Ergebnisse die Integration stoppen — seine eigene Policy, im Vokabular des Gates. Was nicht wandern darf, ist die Definition. Der Test zielt jetzt auf die Tabellenzeile, nicht auf das Wort, und ist dadurch treffsicherer als meine ursprüngliche Absicht.

Convergence Check, erneut und ehrlich

  • Behaviour — spezifiziert und gebrückt für den verifizierbaren Teil. Für den Rest, ob ein Agent Prosa überhaupt befolgt, gibt es in diesem Repository keine Ausführungsschicht: blocked (unavailable), mit [EPIC] Establish Agent Skills conformance and evaluation #65 als dem, was es anderswo beantworten würde. Ausdrücklich keine Waiver — nach der Fassung aus issue_81: Define the canonical Convergence Check #84 ist eine Unmöglichkeit nichts, was Verhältnismäßigkeit akzeptieren könnte, und sie als akzeptierte Abweichung zu verbuchen würde sie begraben.
  • Result: not converged, solange dieser Blocker offen ist. Ob er die Integration verhindert, ist deine Entscheidung, nicht meine — die Evidenz, die hier herstellbar war, ist hergestellt.

./build.sh test 39 JS-Tests / 157 Ruby-Assertions, 0 failures; check-adapters aktuell.

@dieterbaier dieterbaier left a comment

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Vielen Dank für die gründliche Nacharbeit. Der maßgebliche Convergence Check sitzt jetzt am richtigen Zeitpunkt unmittelbar vor der Integration, ein Ergebnis wird an den geprüften Head gebunden, und die ursprüngliche Fehleinschätzung des geänderten Verhaltens wurde korrigiert. Die aktuelle CI ist ebenfalls grün.

Bei den neu hinzugekommenen Prüfungen und beim ausgewiesenen Ergebnis bleiben jedoch drei Punkte:

  1. Der Test für die sieben Fragen prüft nicht die Fragenstruktur.

    Das Szenario behauptet, dass nur der Convergence Check die sieben Fragen enthält. Der Test sucht jedoch lediglich nach der Überschrift Every question must be able to fail. Alle sieben Fragen könnten ohne diese Überschrift in einen aufrufenden Skill kopiert werden, ohne dass der Test fehlschlägt.

    Bitte entweder die tatsächliche siebenstufige Struktur belastbar prüfen oder das Szenario und seinen erklärten Schutzumfang so einschränken, dass sie der implementierten Prüfung entsprechen.

  2. Der Test für die vier Ergebniszustände prüft nur einen Zustand.

    Die Implementierung sucht ausschließlich nach der Tabellenzeile Converged with recorded waivers. Kopierte Definitionen von Converged, Not converged oder Blocked blieben unentdeckt, obwohl das Szenario alle „result states“ abdeckt und die Kommentare behaupten, dass die kanonischen Regeln nicht zurückkopiert werden können.

    Bitte alle vier Definitionen prüfen oder auch hier Szenario und Aussage ehrlich auf den engeren Schutzumfang reduzieren. Da der PR ausdrücklich die Einmaligkeit aller vier Ergebniszustände verifiziert, erscheint die vollständige Prüfung passender.

  3. Finding und Gesamtergebnis des Convergence Checks widersprechen sich.

    Der Kommentar meldet für die Agent-Conformance blocked (unavailable), weist als Gesamtergebnis aber not converged aus. Nach der kanonischen Definition bedeutet erforderliche Evidenz, die an dieser Stelle nicht erzeugt werden kann, Blocked.

    Alternativ – und aus meiner Sicht fachlich sinnvoller – sollte geklärt werden, ob die tatsächliche Befolgung eines Prosa-Skills überhaupt verpflichtende Evidenz für diesen Repository-Change ist. Der Convergence Check verlangt keine vollständige deterministische Beweisführung; er kennt ausdrücklich assisted und human tiers. Wenn die fehlende Ausführungsschicht lediglich ein transparent dokumentiertes Restrisiko mit Follow-up #65 ist, muss sie den PR nicht dauerhaft blockieren. Andernfalls könnte bis zur Umsetzung von #65 praktisch keine Änderung an einem Prosa-Skill konvergieren.

Abgesehen davon ist die erste Nacharbeit überzeugend: Der provisorische und der maßgebliche Lauf sind sauber unterschieden, spätere Commits invalidieren das frühere Ergebnis, und die Integrationsentscheidung bezieht sich auf den tatsächlich zu integrierenden Commit.

Review on #85, three findings, all correct.

The two uniqueness guards claimed more than they checked. One grepped a
single heading, so all seven questions could have been copied without it;
the other checked one of four result states. Both now read what to look
for from the canonical file rather than hard-coding it, so renaming a
state or a question keeps the guard pointed at the right thing instead of
quietly guarding nothing. Mutation probes confirm each fails on the exact
copy it claims to prevent: a pasted "Blocked" definition, and four
question headings without their section heading.

The third finding was a mistake in how I applied the gate, and it exposed
a gap in the gate itself. I reported the missing execution layer for
prose contracts as an unavailable blocker on this change. It is not: it
would appear identically on every change to any prose skill, so blocking
on it would stop all such work until #65 lands while saying nothing about
the change in front of it.

The canonical skill now draws that line. A blocker is evidence this
change owes and cannot produce; a limitation of the medium or the tooling
is a residual risk, which question 6 already requires reporting together
with the follow-up that would remove it. The test is whether the same
blocker would appear on every change of this kind — a gate that can never
be passed is not a gate.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01USXp58FoRppFK6FUht8K7u
@dieterbaier

Copy link
Copy Markdown
Member Author

Alle drei nachgearbeitet in 34f9c0d. Der dritte war kein Formfehler, sondern ein Denkfehler — und er hat eine Lücke im Skill selbst freigelegt.

1 + 2. Die Wächter prüfen jetzt, was sie behaupten

Beide Prüfungen behaupteten mehr, als sie leisteten: eine griff nur eine Überschrift ab, die andere einen von vier Zuständen. Beide lesen jetzt aus der kanonischen Datei, wonach sie suchen, statt es fest zu verdrahten — wird ein Zustand oder eine Frage umbenannt, folgt der Wächter, statt still nichts mehr zu bewachen.

Mutationsproben gegen genau die von dir beschriebenen Lücken:

Mutation Ergebnis
| **Blocked** | … in pr-review eingefügt not ok 3 — skills/pr-review/SKILL.md defines "Blocked"
Vier Fragenüberschriften ohne Abschnittsüberschrift eingefügt not ok 4 — skills/pr-review/SKILL.md carries 4 question headings

Danach aus git wiederhergestellt, Suite wieder grün. Die Szenarien im Feature sind entsprechend umformuliert und behaupten jetzt exakt den Schutzumfang, den der Test hat.

Für die Fragen wird nicht die Abschnittsüberschrift geprüft — genau dein Punkt —, sondern die Titel selbst als Überschriften, ab zwei Treffern. Ein einzelnes geteiltes Wort wie „Traceability" ist kein Kopie-Indiz, zwei sind es.

3. Du hast recht, und der Skill hatte die Lücke

Mein Ergebnis war in sich widersprüchlich — ein Blocker ergibt Blocked, nicht not converged. Aber die eigentliche Korrektur ist deine zweite: das war überhaupt kein Blocker dieser Änderung.

Dein Reductio trifft: Wenn „ein Agent könnte Prosa nicht befolgen" blockiert, kann bis #65 keine einzige Änderung an irgendeinem Prosa-Skill konvergieren. Ein Gate, das nie passierbar ist, ist kein Gate.

Der Fehler war eine fehlende Unterscheidung, die ich beim Schreiben von #84 nicht gesehen habe. Sie steht jetzt im kanonischen Skill:

A standing limitation is not this change's blocker. … The test is one question: would this same blocker appear on every change of this kind? If yes, blocking on it stops all such work indefinitely while changing nothing about the change in front of you.

Fehlende Ausführungsschicht, fehlende Umgebung, noch nicht existierendes Werkzeug — Eigenschaften des Mediums, nicht Lücken der geprüften Änderung. Sie sind Restrisiken, und Frage 6 verlangt ohnehin schon, sie samt Follow-up zu berichten. Dass die Regel aus deinem Review kommt und nicht aus meinem Entwurf, gehört zur Sache dazu.

Convergence Check, korrigiert

  • Result: converged.
  • Behaviour — spezifiziert und gebrückt für das, was hier herstellbar ist: Referenzauflösung, Gate-Erreichbarkeit, Einmaligkeit aller vier Ergebnisdefinitionen und der Fragenstruktur, jeweils mutationsgeprüft.
  • Restrisiko, berichtet nach Frage 6 — ob ein Agent den Prosa-Vertrag tatsächlich befolgt, ist in diesem Repository nicht verifizierbar. Stehende Grenze des Mediums, Follow-up [EPIC] Establish Agent Skills conformance and evaluation #65. Kein Blocker, keine Waiver.
  • Documentation and delivery state — Beschreibung, Kommentare und Diff beschreiben denselben Stand.

./build.sh test 39 JS / 157 Ruby, 0 failures; check-adapters aktuell.

@dieterbaier

Copy link
Copy Markdown
Member Author

Fast geschafft – die technische Nacharbeit passt, aber ein blockierender Punkt bleibt:

Die PR-Beschreibung ist veraltet und widerspricht dem aktuellen Stand.

Sie behauptet weiterhin:

  • Behaviour — not applicable
  • keine Änderung beobachtbaren Verhaltens
  • nur drei Prosa-Verträge, kein ausführbarer Pfad
  • 35 JavaScript-Tests

Der aktuelle PR enthält jedoch zusätzlich Gherkin-Spezifikation und automatisierte Tests, behandelt das Verhalten ausdrücklich als relevant und führt 39 JavaScript-Tests aus. Das neue Convergence-Ergebnis steht nur in einem Kommentar, obwohl der Workflow die Dokumentation im PR-Body verlangt.

Damit ist ausgerechnet Frage 7 des Convergence Checks noch nicht erfüllt: Beschreibung, Kommentare und Diff erzählen nicht dieselbe Geschichte. Die PR-Beschreibung sollte auf den aktuellen Head 34f9c0d aktualisiert und das maßgebliche Ergebnis dort festgehalten werden.

Ansonsten habe ich keine weiteren Findings:

  • Beide Wächter lesen ihre Vergleichswerte jetzt aus dem kanonischen Skill.
  • Alle vier Ergebnisdefinitionen werden geprüft.
  • Die tatsächlichen Fragentitel werden geprüft und das Szenario benennt den Schwellenwert ehrlich.
  • Die neue Abgrenzung einer stehenden Einschränkung vom change-spezifischen Blocker ist nachvollziehbar.
  • CI Validate Run generate the refinement note instead of hand-writing it #68 ist grün.

Nach Aktualisierung der PR-Beschreibung wäre der PR aus meiner Sicht merge-bereit.

@dieterbaier

Copy link
Copy Markdown
Member Author

PR-Beschreibung auf Head 34f9c0d aktualisiert, mit dem maßgeblichen Convergence-Ergebnis im Body statt nur im Kommentar — der Workflow, den dieser PR selbst schreibt, verlangt es dort.

Entfernt: Behaviour — not applicable, "kein ausführbarer Pfad", die drei Prosa-Verträge, 35 Tests. Ergänzt: die Gherkin-Spezifikation mit ihren vier mutationsgeprüften Wächtern, 39 Tests, und die Regel, die aus diesem Review in den kanonischen Skill gewandert ist.

Bemerkenswert ist, welche Frage es war: Frage 7, zum zweiten Mal in dieser Kette. Auf #83 war es die Skill-Zahl, hier der ganze Verhaltensabschnitt — beide Male ist der Code konvergiert und die Delivery-Metadaten sind nicht nachgezogen, beide Male hat ein Mensch es gefunden und nicht ich. Zwei Instanzen desselben Musters in drei PRs sind kein Zufall, sondern ein Hinweis darauf, wo diese Arbeitsweise systematisch driftet.

Merge jetzt.

@dieterbaier
dieterbaier merged commit f6feb53 into main Aug 31, 2026
1 check passed
dieterbaier pushed a commit that referenced this pull request Aug 31, 2026
Review on #85 found two connected problems.

The check ran too early. The step sat before "declaring the pull request
ready", while the canonical skill puts the authoritative run before
mergeability and says plainly that an early pass is not a result. Review
comments and later commits change the state it inspected, and PR
Integration only read the recorded result without asking whether it was
still current.

The run at PR-open time is now explicitly provisional. Address PR
Comments voids a recorded result when new commits land, and PR
Integration runs the authoritative check against the commit that would be
integrated. A result that does not cover the current head joins the
warning signs, because a stale result is not a result.

The second finding was mine to own: this change alters observable agent
behaviour — agents must run and record the gate, and must not integrate
on certain results — so reporting Behaviour as "not applicable" was
wrong. The skills are the behaviour contracts here.

What can be specified now is, and it is bridged: skill-to-skill
references must resolve, the three callers must reach the gate, and its
result-state definitions and question structure must live in exactly one
file. A mutation probe confirms the guard fails when a reference is
broken.

Writing that guard sharpened a distinction the prose had left implicit.
Naming a result state elsewhere is legitimate — implement-issue-workflow
says which results stop integration, which is its own policy in the
gate's vocabulary. Defining one elsewhere is the copy that drifts, so the
test targets the definition, not the word.

What remains unverifiable here is whether an agent obeys prose at all.
That is #65, and it is reported as an unavailable blocker rather than
waived.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01USXp58FoRppFK6FUht8K7u
@dieterbaier
dieterbaier deleted the issue_82 branch August 31, 2026 14:47
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Slice #77.3: Invoke the Convergence Check from the participating skills

1 participant