docs(memory): a permission-mode verdict is not a memory (CLOUD-178, CLOUD-1166) - #791
docs(memory): a permission-mode verdict is not a memory (CLOUD-178, CLOUD-1166)#791wenzowski wants to merge 1 commit into
Conversation
CLOUD-178 claude.ai connector tools flip between readable and UUID names, silently breaking the permission allowlist
Split out of CLOUD-177, which is Done on its own scope. Evidence thread: the comments on CLOUD-177, in particular the final correction establishing verdict (a). ReadyA claude.ai connector MCP server is exposed to a session under two different names over its lifetime, and a permission allow rule can only name one of them. At session start the Linear connector appears as This hits the board-in-lockstep rule directly: moving a Measured, across two containers
Verdict: the UUID is stable and equal across containers. It is a durable identity for the connector; what varies is which of the two names is live at a given moment. This was initially mis-called as "unstable per boot" from a session that had not yet reconnected — absence of a reconnect is not evidence of naming stability, and any future measurement here needs to span at least one reconnect to say anything. Confirmed load-bearing, not cosmetic: a Note this is a second, independent defect from the one that originally masked it: an org-level Ready predicateA session survives a connector reconnect with Linear calls still auto-approved — no prompt on a Done (proposed)
"mcp__Linear",
"mcp__4db58e41-cd4e-4818-8922-46cf616593f4"
Open questions
A fourth connector, measured 2026-08-19, and it is NOT LinearThe injected config in one remote session ( The symptom that surfaced it: This is the first instance where the affected connector is the one the fleet procedure runs on.
Mitigation applied, user-level only (repo rule 1, and the Done section's reasoning above): One trap worth recording, because the obvious form of this mitigation is unsafe. A wildcard grant ( A second open question is now partly answered. The readable-phase test PR #488 called for has now run, and it FAILED — so the flip is not the only causePR #488 closed with an explicit outstanding test: "This session is in the UUID-exposure phase, where no readable-name rule matches either, so the fix cannot be exercised here. The next session that starts in the readable phase is the test." That session has happened. It started in the readable phase, and the corrected spelling did not stop the prompting.
Name, spelling and rule matched exactly, and the call prompted anyway. So the two-name flip is not sufficient to explain the prompting, and neither is the retired typo theory. Something above The candidate, recorded as a hypothesis and NOT as a finding. This session ran What this bounds. CLOUD-191's resolver — built today — repairs the UUID episode, and that repair is real and independently worth having, because during such an episode the two committed denies enforce nothing and AGENTS.md's ban on babysitting timers rests on them. But it cannot repair the readable phase, because in that phase there is no name mismatch to translate. A session that lands CLOUD-191 and still sees prompts under the readable name has not found a regression — it has found this, and should add its observation here rather than re-diagnosing the flip. The CLI's own MCP logs carry a STABLE identity, and they date the flip inside a single session (measured 2026-08-31)Every previous pass at this row worked from the tool NAME, which is the one identifier that flips. Three identifiers, and they do three different things. Read on one machine across 2026-08-28 → 2026-08-31:
The two UUIDs were never distinguished before, and conflating them is why "is the UUID stable?" kept resolving differently. One is stable and one is not, so the answer depends entirely on which was read. This answers a stated open question and retires a mitigation this row calls unsafeThe row records, of the remote-session grants: "which UUID is
Both halves are corroborated independently of the log: this session's own tool list carries Miro is a FIFTH connector this row has never recorded, and it is the cleanest case: readable-name directory with exactly one file, from session start, and nothing since. The flip is dated, and it happened INSIDE one sessionEvery log line carries
Every readable name stops on 2026-08-28; every UUID continues. Both start at the same instant, so the two names were live in the same session batch and the readable one was then dropped and never came back. That is the flip, with a timestamp, from the host's own record rather than from a human noticing a prompt. What this does NOT establish, stated so the next reader does not over-read it
The local OAuth store is empty, so auth rides the injected headers
The classifier paragraph above is a hypothesis and a memory had promoted it to a ruleThis row is careful: it labels the auto-mode classifier a "hypothesis and NOT a finding" and names the unrun experiment ( Why the repo cannot fix this itselfIn Claude Code on the web, connectors are "provisioned by the remote host and arrive as explicit Generated by Claude Code CLOUD-1166 A COUNT — or a membership LIST — copied out of another row's body is republished as “today” and nothing re-derives it: 37/49 against a tree holding 52/66, three mutated copies of one blocked-class list, and two blocked verdicts asserted from a bucket nobody opened
The instance. CLOUD-1117's body carried "37 verdict rows and 49 routes" as a statement about the current tree. A decision record written against that row on 2026-08-29 restated the pair, in bold, twice, with the word today. Measured the same day on Both citations were wrong, by 15 rows and 17 routes. The registry had grown ~40% since the number was written, and neither the source row nor the citing record could tell. The class, stated as a predicate rather than as this instance. A quantitative claim about the tree can be written into a row's body, where it is a measurement with a timestamp, and then cited by another row, where it silently becomes a premise. Every mechanism this repository has stops at the first hop:
This is CLOUD-686's shape over a number instead of a condition. CLOUD-686 is the deferral whose reversal condition is later satisfied while nothing re-fires. A cited count is the same defect with the condition implicit: the row asserts a state of the tree, the tree moves, and the assertion is load-bearing in a second row that never looked. CLOUD-929 is the same failure with the census in the body rather than cited from it; CLOUD-633 is it over a prose literal's recall. Three instances, no mechanism. Why the citing session did not catch it here, which is the honest part. That session read §1 of CLOUD-1117, which named Refinement — Ready Source of truth (§1). The tree is the only authority for a claim about the tree. Computable deliverable (§2). A cited-count notation with a gate behind it, so a number carries the command that produced it. Concretely: a body may state a tree count only in a form that names its own reproducing command, e.g. and a gate re-runs the cited command and compares. Two arms, and the second is what makes it a gate rather than a convention:
The command must resolve to a value, not to prose — non-negotiable rule 3 — so the notation binds a single-integer-producing command and nothing else. The pointer emitted is the count and Acceptance predicate (§2).
Effect (§3). A new gate task reading an issue payload on stdin, like its siblings. It executes the cited command, which is an effect: the command set must be bounded — the read-only allowlist (house style §5) is the existing boundary and the gate must not become a general shell evaluator over issue bodies. That bound is the hard part of this row, not the comparison, and a design that skips it is not landable. Generated artifacts and drift (§4). None expected. If the notation is parsed in Rust rather than in bash, Output and exit contract (§5). Exit 1 for a stale or uncited count; exit 2 where the cited command could not be run at all — could-not-look, never a silent pass, and never a refusal that blames the author for the gate's own inability to look. **Commit / bump (§6). ** Test and replay obligation (§7). A **Relationships (§8). **CLOUD-686's shape over a number; CLOUD-929 and CLOUD-633 are the two prior instances. Found while remediating CLOUD-1117. |
|
Warning Review limit reachedNext included review available in 38 minutes. View limit detailsLimit details: You’ve used the included review currently available. Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available. Review configuration: ⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Free Run ID: 📒 Files selected for processing (3)
Note 🎁 Summarized by CodeRabbit FreeYour organization is on the Free plan. CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please upgrade your subscription to CodeRabbit Essentials by visiting https://app.coderabbit.ai/settings/billing. Comment |
Two bullets recorded a session-scoped permission reading as a durable
capability boundary, and both were later quoted to refuse work:
- connector-allowlist-recovery said replaying the injected config's
headers was "blocked by the auto-mode classifier, correctly ... Do not
route around it". CLOUD-178 is careful to label that classifier a
hypothesis and names the experiment nobody has run; the memory promoted
it to an imperative. The durable half stays: what the headers ARE
(X-MCP-Server-ID, X-MCP-Server-Origin, X-Session-UUID, no bearer), and
that they are session-bound and agent-readable.
- prior-art-and-issue-hygiene said `add_repo` on a survey target "is
declined by the permission classifier", which contradicts the standing
instruction to call it rather than pre-judge -- an unauthenticated probe
of a private repo 404s whether or not the session is authorised, so a
pre-check reports a false negative.
workflow/agent-fanout keeps its permission-mode paragraph: it records a
divergence between what create_session's schema documents ("omit to
inherit it") and what was measured, which is what a memory is for.
core.md, the graph root, now carries the rule: a memory records what is
re-derivable -- a path, a file format, an argv, a measured flip, an
upstream issue number. If a claim would change depending on which
permission mode the session is in, it is not a memory. The tell is an
imperative with no re-derivable object behind it.
Refs: CLOUD-178, CLOUD-1166
45dde19 to
242b883
Compare
|
❌ The last analysis has failed. |
Refs: CLOUD-178, CLOUD-1166. Closes neither — both stay open on their own scope.
What this changes
Two memory bullets recorded a session-scoped permission reading as a durable capability boundary, and both were later quoted back to refuse work. This excises them, keeps the re-derivable half of each, and states the rule in
mem:coreso the class does not recur.connector-allowlist-recovery.mdRemoved: "Replaying the injected config's
headers… is blocked by the auto-mode classifier, correctly — it is credential replay. Do not route around it."CLOUD-178 is careful about exactly this claim — it labels the auto-mode classifier "a hypothesis and NOT a finding" and names the experiment nobody has run (
autovsdefault, same tree, same verb). The memory promoted it to an imperative, and it was then cited to refuse a design that was not credential replay.Kept, because it is re-derivable and load-bearing: what the headers are —
X-MCP-Server-ID,X-MCP-Server-Origin,X-Session-UUID, no bearer token, so the endpoint authenticates the session — and the two consequences that follow regardless of any session's permission state (the material is session-bound and cannot be carried; it sits in an agent-readable file).prior-art-and-issue-hygiene.mdRemoved: "
add_repoon a survey target is declined by the permission classifier."This contradicts the standing instruction to call
add_reporather than pre-judge it — an unauthenticated probe of a private repo returns 404 whether or not the session is authorised, so a pre-check reports a false negative. Rewritten to say the tool's own answer is the authority.workflow/agent-fanout.md— deliberately NOT changedIts permission-mode paragraph stays. It records a divergence between what
create_session's schema documents ("omit to inherit it") and what was measured, which is precisely what a memory is for. Naming the one that stays is the point: the rule below is not "delete anything mentioning permissions."core.md— the rule, at the graph rootWith the tell: an imperative with no re-derivable object behind it ("do not route around it", "is declined by"), where the honest form is the attempt and the tool's own answer.
Why the graph root
mem:coreis where every other memory's trigger already lives, so a rule about what may be written to any of them belongs there rather than duplicated per file. It sits immediately under the "read on demand, never all of them" paragraph.Verification
hkpre-commit gate green:rules-drift,module-map-check,suite-bench-check,license-table-check,no-docs-tree,prettier,conventional-commit,commit-attribution.edit_memory—.serena/memories/**isprotected, so a shell write is refused by the engine's protected-path gate.Branch note
Two earlier commits on this branch were already on
main, and a third (963bff95) was a duplicate of landed work whose only remaining delta was stale counts —maincarries the preset arm atpolicy/shell-retirement.rego:946with its test, and newer ledger figures (725 arms / 113 engine-source / 21 whole-file) than that commit's 609 / 110 / 18. The branch was reset onto currentorigin/mainand only the new commit replayed, so the stale counts do not travel.Generated by Claude Code