Skip to content

runner: a live lease runs the shift, and the nightshift skill is code now - #18

Merged
JulienMartel merged 1 commit into
mainfrom
worktree-lease-runs-the-shift
Sep 6, 2026
Merged

JulienMartel merged 1 commit into
mainfrom
worktree-lease-runs-the-shift

Conversation

@JulienMartel

Copy link
Copy Markdown
Contributor

A live lease runs the shift on its own. factory lease grant 12h starts factory watchdog run, and that process is now the runner: it passes factory shift --json every 20 minutes, puts every red default branch through the four fixer gates, and revokes a lease whose passes stopped landing. No agent pane has to survive the night for a docs PR to merge at 3 a.m.

The nightshift skill is deleted. Its judgement was four string checks and a retry counter, and a rule that is four string checks is code: deterministic, tested, and reachable by a case that fails when the gate is deleted.

What changed

The runner (libexec/factory-watchdog)

  • One pass every runner.interval (1200s). The first lands within seconds of the grant.
  • A pass that could not see (prs-unknown, tier-unknown, ci-unknown, after-merge-failed, pass ABORTED) gets one more pass at the next tick, once. pass-retry says so.
  • The four fixer gates, in the order their reason is quoted: fixer.command configured; no fixer-spawned: <repo> <sha> in any shift log; fewer than fixer.cap (2) lanes for the repo in the day's log; the pass's budget event said fixer: true. Every outcome is a line: fixer-spawned, fixer-skipped naming the gate, or fixer-failed with the command's stderr and a card.
  • fixer.command is handed <repo> <default branch> <run url>. Default none, so a machine without haus opens nothing.
  • A shift that exits before writing anything is pass-failed, and not a heartbeat. The stale/dead thresholds keep their meaning under a runner: shift-stalled at 45 minutes, shift-dead and a revoke at 90. shift-over closes a timed lease's log.
  • FACTORY_RUNNER_INTERVAL joins the three env overrides that may only shorten. watchdog.stale has to be greater than runner.interval, checked at load and again over the overrides.

The lease (libexec/factory-lease)

  • factory lease grant indefinitely. The state file spells it never, a word rather than a sentinel epoch, so a reader that only knows epochs fails closed. status --json carries indefinite: true with expires and secondsLeft null.

The shift (libexec/factory-shift)

  • The ci-red event carries head and branch. The metered budget event carries reason on refusal. The runner reads events, never the human line.

The surface

  • config print has a runner row and a fixer row. doctor checks fixer.command the way it checks notify.command: not on PATH blocks, none configured is a note.
  • watchdog once says NO RUNNER at exit 4; its JSON field is runnerPid.
  • ai/nightshift/SKILL.md is gone. The factory skill says what the runner does and what an agent no longer does (loop the shift, spawn a lane by hand).
  • README: the four verbs, the lease section, a new section on the runner replacing "When the foreman dies", and "Driving it from an agent" rewritten. AGENTS.md carries the gate rule.

Verified

bats test/ is green at 242 cases, run three times over the watchdog suite to shake out timing. shellcheck -x over every shipped script is clean, and script/check-skills.sh passes with one skill.

New cases: a live lease and the runner produce pass done lines with no agent involved; a CI-RED spawns through fixer.command with the three words in order; one case per gate, each written so that deleting the gate fails it; the retry is once and at the next tick; a shift that cannot start reaches shift-dead; an indefinite lease grants, reads back, revokes, and is revoked by shift-dead too.

Landing order

Merge this PR, but do not bench ship factory until the haus slice has merged: haus names nightshift in modules/ai/tool-skills.nix, and a lock bump to a factory without it fails every haus rebuild. Ship both in one ripple.

Things: GPvZXPXg8xypNB7V3piSo7 under the project night shift as a daemon.

🤖 Generated with Claude Code

https://claude.ai/code/session_013cfzqFVmVMNPYnVQteAhxn

… now

`factory lease grant` starts `factory watchdog run`, and that process is the
runner: it passes `factory shift --json` every runner.interval (1200s) while
the lease is live, retries a pass that could not see once at the next tick,
puts every ci-red through the four fixer gates, and revokes a lease whose
passes stopped landing. No agent pane has to survive the night.

The four gates were the nightshift skill's prose — four string checks and a
retry counter, which is code: fixer.command configured; no fixer-spawned for
the same head SHA in any shift log; fewer than fixer.cap (2) lanes for the
repo in the day's log; the pass's budget event said fixer: true. Every
outcome is a line — fixer-spawned, fixer-skipped naming the gate, or
fixer-failed with stderr and a card — and every gate has a case that fails
when it is deleted. fixer.command is handed <repo> <default branch> <run url>
and defaults to none, so a machine without haus opens nothing.

`factory lease grant indefinitely` writes `never` where the epoch goes — a
word, so a reader that only knows epochs fails closed — and status --json
carries indefinite: true with expires and secondsLeft null.

The shift's ci-red event carries head and branch, and the budget event its
reason, so the runner reads events and never a human line. `watchdog once`
says NO RUNNER and its JSON field is runnerPid; the log lines are
shift-stalled / shift-resumed / shift-dead / shift-over / pass-failed /
pass-retry. doctor checks fixer.command the way it checks notify.command,
config print has a runner row and a fixer row, and watchdog.stale must sit
above runner.interval, at load and over the env overrides.

ai/nightshift/SKILL.md is deleted; the factory skill carries what an agent
still does and what it no longer does. README: the lease section, a new
section on the runner in place of "When the foreman dies", and "Driving it
from an agent" rewritten. Not shipped to haus yet — its tool-skills.nix
names the deleted skill, so the two land in one ripple.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013cfzqFVmVMNPYnVQteAhxn
@JulienMartel
JulienMartel merged commit 38d5a01 into main Sep 6, 2026
2 checks passed
JulienMartel added a commit to hausfold/haus that referenced this pull request Sep 7, 2026
The layer, not the desktop: a new `haus.ai.factory.enable` (default follows
`haus.ai.enable`) and the per-user launchd agent `com.hausfold.factory` behind
it, running `factory watchdog run`.

factory's runner is code as of hausfold/factory#18: a live lease runs `factory
shift` on a cadence and puts every red default branch through four gates that
used to be a skill an agent session re-read on every wakeup. That session was
the foreman, and factory's watchdog watched it. With the judgement in code the
only supervision left is restarting a runner that died, and `KeepAlive` is that
supervisor — a reboot, a panic or an OOM kill is now a restart instead of a
lease standing all night with nobody exercising it.

`ThrottleInterval = 300` is load-bearing and measured. `run` exits 0 in ~0.2s
with no live lease, which is this machine almost always, and the default
10-second throttle would make that ~8,600 spawns a day. It costs the crash case
nothing because launchd measures the window from the last SPAWN, not the exit:
a probe at `ThrottleInterval = 20` killed 8s in waited the remaining 12s, and
killed after 30s alive came back in under a second. A runner under a live lease
has been up at least one `runner.interval` before anything can kill it.

`SuccessfulExit = false` is the shape that looks right and is not: it would
leave the job down after every lease-less exit, so a `factory lease grant` typed
in another pane or over ssh would be supervised by nothing.

`haus-factory-fixer` is the second half. factory's runner appends `<repo>
<default branch> <run url>` to `fixer.command`; `haus-fix-github` takes
`<selector> <verdict> <url>`, and `ci` is carried by neither side. That mismatch
was written down in docs/night-shift-internals.md as a shim the reader had to
write; the layer ships it now, refusing any argv that is not exactly three words
so a bad `fixer.command` cards a `fixer-failed` instead of opening a lane on a
branch nobody named. The policy file stays the person's — it is authority, and
haus writes none of it.

Verified: `nix flake check` green, `bench try` builds, the plist and the shim
read back correct from the store, and the throttle semantics above were measured
with a probe job rather than reasoned about.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011WXH88Rggk5hif6x32hgkp
JulienMartel added a commit to hausfold/haus that referenced this pull request Sep 7, 2026
* ai: launchd owns factory's runner, and the fixer argv shim ships

The layer, not the desktop: a new `haus.ai.factory.enable` (default follows
`haus.ai.enable`) and the per-user launchd agent `com.hausfold.factory` behind
it, running `factory watchdog run`.

factory's runner is code as of hausfold/factory#18: a live lease runs `factory
shift` on a cadence and puts every red default branch through four gates that
used to be a skill an agent session re-read on every wakeup. That session was
the foreman, and factory's watchdog watched it. With the judgement in code the
only supervision left is restarting a runner that died, and `KeepAlive` is that
supervisor — a reboot, a panic or an OOM kill is now a restart instead of a
lease standing all night with nobody exercising it.

`ThrottleInterval = 300` is load-bearing and measured. `run` exits 0 in ~0.2s
with no live lease, which is this machine almost always, and the default
10-second throttle would make that ~8,600 spawns a day. It costs the crash case
nothing because launchd measures the window from the last SPAWN, not the exit:
a probe at `ThrottleInterval = 20` killed 8s in waited the remaining 12s, and
killed after 30s alive came back in under a second. A runner under a live lease
has been up at least one `runner.interval` before anything can kill it.

`SuccessfulExit = false` is the shape that looks right and is not: it would
leave the job down after every lease-less exit, so a `factory lease grant` typed
in another pane or over ssh would be supervised by nothing.

`haus-factory-fixer` is the second half. factory's runner appends `<repo>
<default branch> <run url>` to `fixer.command`; `haus-fix-github` takes
`<selector> <verdict> <url>`, and `ci` is carried by neither side. That mismatch
was written down in docs/night-shift-internals.md as a shim the reader had to
write; the layer ships it now, refusing any argv that is not exactly three words
so a bad `fixer.command` cards a `fixer-failed` instead of opening a lane on a
branch nobody named. The policy file stays the person's — it is authority, and
haus writes none of it.

Verified: `nix flake check` green, `bench try` builds, the plist and the shim
read back correct from the store, and the throttle semantics above were measured
with a probe job rather than reasoned about.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011WXH88Rggk5hif6x32hgkp

* ai: name the launchd state the services deck's honesty rests on

`KeepAlive = true` classes this job `running`, and doctor calls a `running` job
that is not live wedged — which with no lease would be a permanent red line for
the majority state. It is not, and the reason is a launchd state rather than
anything in the entry: a KeepAlive job waiting out its throttle prints `state =
spawn scheduled`, never `not running`, and `_svc_probe` reads only the latter as
stopped. Sampled every 30s across a full 300s window on a probe job: `spawn
scheduled` throughout, `runs` ticking at the boundary.

Written down in both places because the next person to touch `ThrottleInterval`
or that state test needs to re-check the pair.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011WXH88Rggk5hif6x32hgkp

* ai/core: the assurance pass's four, and a build that holds the fixer argv

From the pre-PR review of this branch. Nothing was ≥3/5; these are the cheap
ones worth doing rather than carrying.

`factory-fixer-argv`, a new `nix flake check`: `haus-factory-fixer` is a pure
argv reorder and both ends of it live outside its file — factory's runner
appends `<repo> <branch> <run url>`, `haus-fix-github` reads `<selector>
<verdict> <url>`. Neither side has a reason to tell haus it moved and the
failure is silent: a lane briefed at 3 a.m. on a branch name parsed as a repo,
with a `fixer-spawned` line saying the night went fine. `factory doctor` cannot
see it either — it blocks on a `fixer.command` PATH cannot find and checks
nothing about its argv. So the check reads all three, the locked factory
included, exactly as `pounce-item-grammar` reads the locked pounce.
Negative-tested: flipping the shim's `$2` to `$1` turns it red with the right
message.

The runner's PATH gains `/nix/var/nix/profiles/default/bin`. Determinate owns
the nix daemon, so `nix` is there and nothing else in that PATH names it. Both
after-merge hooks this family runs self-repair, but `afterMerge.commands` is
arbitrary shell a person wrote, and one calling `nix` directly would die
overnight with the only trace in /tmp/haus-factory.err.log. Renamed to
`userPath` while there: `grep userPath modules` is how somebody finds every
launchd PATH in this repo, and a private name took this one out of that answer.

The deck entry no longer says "every twenty minutes" — `runner.interval` is a
factory config value haus cannot read, and the deck's own rule is never to copy
a fact you do not own.

And the launchd job counts: the repo declares eighteen now, six gated with
`lib.mkIf` on their own line and twelve expecting to be up. Ten call sites in
AGENTS.md, modules/core/default.nix and modules/core/haus.sh, the same sweep
1d7bf88 did for sixteen → seventeen. The three historical ones ("until this
deck existed…") keep their old numbers, which is what makes them history.

Verified: `nix flake check` 47 green, `bench try` builds.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011WXH88Rggk5hif6x32hgkp

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant