Skip to content

Auto Watch: time other mods' patches on hot methods within the 40-method cap (stacked on #2) - #5

Merged
kramsey458 merged 4 commits into
mainfrom
claude/auto-watch
Sep 22, 2026
Merged

kramsey458 merged 4 commits into
mainfrom
claude/auto-watch

Conversation

@kramsey458

Copy link
Copy Markdown
Contributor

Stacked on #2 (#2, base branch claude/sampling-calibration). It builds on #2's per-method sampling, so merge #2 first. After that, retarget this PR to main.

Summary

New setting AutoWatch (only in PerformanceLog.cfg, off by default). When it is on, the mod times the prefixes, postfixes and finalizers that other mods put on the game's hot methods. It uses the Watch slots the Watch entries leave free (40 methods in all). This is the "PerformanceLog auto Watch" stretch item.

Today a mod's prefix on every entity's tick (BeaverBuddies) or on a singleton's Tick (Late Game Performance) runs inside a profile row that names the game or the singleton, so its cost never shows up anywhere. With AutoWatch = true, each of those patch methods gets its own method row in profile.csv, tagged with the patching mod. The frames.csv header also gets a # watch|...|auto|<kind> on <method>|<owner> line for each one.

The mod only observes. It never unpatches, reorders or changes another mod's patch, and never calls UnpatchAll. The frame path is unchanged and allocates nothing (a check measures this). The auto watch also puts the game's random state back after it makes its patches, so BeaverBuddies co-op stays in step (see Review).

Items

AUTOWATCH (brief section 6, "PerformanceLog auto Watch")

What changed

  • Choosing the methods. source/Core/AutoWatch.cs (AutoWatch.Plan) is pure code with no Unity, Timberborn or Harmony types.
    • It takes patches on the methods behind the profile's own rows first: a singleton's Tick, UpdateSingleton, LateUpdateSingleton or StartParallelTick, or an entity's or component's Tick. Then it takes the rest of PatchFormat's hot list.
    • Within each group the order is by name: patched method, then prefix, postfix, finalizer, then patch method. So the choice is the same whatever order Harmony lists the patches in.
    • A patch method on several hot methods counts as one watch, placed by the first of them.
    • Watch entries win: they take their slots first, and a method a Watch entry already watches is not watched twice.
    • These take no slot: a refused method (abstract, generic, no body, or a catch ... when filter) and a patch that fails.
    • This mod's own patches and transpilers are never taken.
    • A prefix that returns bool can skip the method it patches. Its line says (can replace it).
  • Making the patches. source/Game/Watch.cs (AutoInstall, Examine, RunsBehindAProfileRow, AutoCapability, AutoLines, AutoFinal) reads Harmony's registry and makes the patches once per run of the game, when the first log starts (Session.Start, at PostLoad, before the first tick). All of it runs inside Watch.KeepUnityRandom, which saves UnityEngine.Random.state and restores it in a finally.
  • Output.
    • # capability|autoWatch| at the start and # capability-final|autoWatch|N of M ... were called at the end. The final line also lists the methods never seen called: not called, called only off the game thread, or inlined by the runtime.
    • One # watch| line per patch method, including the ones left out and why.
    • One capability line in summary.md.
  • Analyzer. perflog.py and summary.md name method rows Class.Method(...), e.g. TickableEntityTickPatcher.Prefix(TickableEntity), instead of a bare Prefix(...). The report adds auto watch: <kind> on <target> (<owner>) under each auto-watched row. compare no longer counts a method that only one side watched as extra time.
  • Docs. README (settings table, an AutoWatch section with the co-op paragraph, "Working with other mods"), SESSION-README ("Which mod is it?"), CLAUDE.md (layout; the one patching done after StartMod; the BeaverBuddies RNG rule), TESTING.md (counts, table row, in-game items 9 and 10), packaging cfg, and the in-game Mod Settings note.

Tests that prove it

tests/AutoWatchTests.cs has 7 checks:

  • CapIsRespected
  • WatchEntriesWin
  • OrderIsByName: 20 shuffles, plus the tie-break for a patch method on several hot methods, and the replace mark.
  • PatchesAreExamined: the real TickableEntity.Tick, MeteredTickableComponent.Tick, Walker.Tick and GameSaver.Save, and a bool prefix.
  • WatchedPatchIsTimed: 2000 calls allocate 0 bytes; also checks the header lines and the never-called list.
  • EntriesKeepTheirSlots: Watch entries' slots and names, through AutoInstall.
  • GameRandomIsKept:
    • The default guard is KeepUnityRandom.
    • Its IL reads UnityEngine.Random.state before its try and writes it back in its finally.
    • Every patch, and the registry read, happens inside the guard.

Other checks:

  • GameBindingTests ConfigParsing, ConfigProblems and SettingsApplyTo cover the key.
  • SummaryTests.ShortNames covers Summary.ShortMethod.
  • tools/test_perflog.py has 3 cases: the report's auto-watch lines, compare with AutoWatch off versus on, and the report's line when the auto watch produced no rows.

Fail before / pass after

  • First commit (9ed9f2e). The tests were written against stub signatures and failed on their assertions: C# 97/104 passed, e.g. expected 9, got 0. Python failed on 'TickableEntityTickPatcher.Prefix(TickableEntity)' not found. After the commit: C# 104/104, Python 47 OK. An independent reviewer reverted the source and confirmed the tests fail without it: 96/104 with behaviour-free stubs.

  • Review commit (6ea1498). Mutation run (_plE_fin_mutants.log and _plE_fin_mutants2.log). Every mutant made the suite fail (exit 1):

    Mutant Check that failed
    Guard does not restore the state GameRandomIsKept
    Patching bypasses the guard GameRandomIsKept
    Default guard is a no-op GameRandomIsKept
    Cap ignores the Watch entries EntriesKeepTheirSlots
    Watch entries' names ignored EntriesKeepTheirSlots
    Never-called list dropped WatchedPatchIsTimed
    First-seen wins in the tie-break OrderIsByName
    Replace mark dropped OrderIsByName
    summary.md shortens methods like other names SummaryTests.ShortNames

    The two new Python cases fail against the first commit's perflog.py. That run printed the exact wrong sentence the reviewer found: Only B has these ... TickableEntityTickPatcher.Prefix(TickableEntity) (2.0 ms/s). Time B spends that A does not.

Tests

Suite Base claude/sampling-calibration 2e4d468 After 9ed9f2e Head 6ea1498
C#: dotnet run --project tests -c Release -p:GameDir=... -- --managed ... 99/99 104/104 106/106
Python: python -W error::ResourceWarning -m unittest discover -s tools -p test_perflog.py 46 OK 47 OK 49 OK

The mutation runs showed a known flake twice: WriterTests "a file someone holds without sharing is retried" and "a file can be read while the writer runs" each failed once, both with frames.csv ... being used by another process. Many agents were building at once. Neither failed in a normal run, and neither is related to this change.

Co-op impact

  • No simulation or wire-format change. It only observes, and it is off by default.
  • The auto watch patches after BeaverBuddies has seeded the game's random numbers, and under BeaverBuddies every Harmony patch draws from them (see Review). It therefore restores UnityEngine.Random.state after its patches, so the game should play out exactly as without it. That is reasoned from the code and a unit check of the guard's IL. It has not been seen in a two-player game (TESTING.md item 10).
  • Until item 10 has passed, keep AutoWatch = false (the default) for every co-op player.

Proposed CHANGELOG lines

  • New setting AutoWatch (PerformanceLog.cfg only, off by default). It times the patch methods other mods put on the game's hot methods, in the Watch slots the Watch entries leave (40 methods in all; Watch entries always come first).
    • It takes patches on the per-tick and per-frame methods behind the profile's rows first (a singleton's Tick/UpdateSingleton/LateUpdateSingleton/StartParallelTick, or an entity's or component's Tick), then the rest of the hot methods, in name order.
    • Each gets a method row in profile.csv and a # watch|...|auto|<kind> on <method>|<owner> line in the frames.csv header. A prefix that can skip the game's method is marked (can replace it).
    • # capability|autoWatch| says what it took and what it left out; # capability-final|autoWatch| says how many were seen called.
    • It adds only its own timing patches, when the first game is loaded, changes no other mod's patch, and puts the game's random state back afterwards (BeaverBuddies co-op).
  • perflog.py and summary.md name watched methods by class and method (TickableEntityTickPatcher.Prefix(TickableEntity), not just Prefix(TickableEntity)). The report says which hot method each auto-watched one is on and whose patch it is. compare no longer counts a method only one side watched as extra time.
  • (Observe-only: no simulation or wire change.)

Needs in-game / two-player testing

  • TESTING.md item 9, AutoWatch smoke test.

    1. Set AutoWatch = true and restart.
    2. Load a save with BeaverBuddies (MultiColony) and Late Game Performance, and play 3 minutes, some of it at speed 3 or more.
    3. Check Player.log: it should say Auto watch: watching N of M patch methods... with N > 0 and no warning.
    4. Check the header: # capability|autoWatch|watching ..., and a # watch|...|watching|auto|... line per method.
    5. Check profile.csv: method rows for them. BeaverBuddies' TickableEntityTickPatcher.Prefix/Postfix should be among them, and busy.
    6. Check that # capability-final|autoWatch| says most of them were called.

    This also verifies, for the first time, the __originalMethod key the config Watch relies on.

  • AutoWatch cost: play the same save for the same time at the same speed with AutoWatch = false, then true. Compare overheadUs/probeUs and the frame rate. The report warns above 2% of a frame.

  • TESTING.md item 10, co-op. Two players connected with BeaverBuddies for 10 minutes or more, with AutoWatch = true for one player only, some of it at speed 3 or more, including a save. There must be no desync, and BeaverBuddies' behaviour and Player.log must be the same as without it.

  • With it on, BeaverBuddies' and Late Game Performance's own behaviour is unchanged: no new errors in Player.log, and saves still work.

Review

Three adversarial reviewers looked at 9ed9f2e.

Fixed

  • Blocker (determinism): AutoWatch drew from the game's RNG after BeaverBuddies had seeded it, a co-op desync.
    • I checked the chain in the BeaverBuddies source. DeterminismService's constructor calls UnityEngine.Random.InitState(seed). GuidPatcher.Prefix on Guid.NewGuid calls GenerateWithUnityRandom(), 16 × UnityEngine.Random.Range, with no load guard. Every Harmony.Patch builds a MonoMod DynamicMethodDefinition, whose field initializer calls Guid.NewGuid().
    • Fix: every patch the auto watch makes, and its reading of the registry, now runs inside Watch.KeepUnityRandom, which restores UnityEngine.Random.state in a finally. If the state cannot be read, nothing is patched. GameRandomIsKept checks both the wiring and the guard's IL.
    • CLAUDE.md now has the rule: any patch made after a game starts loading must run in that guard, and nothing may be patched mid-game. README's co-op sentence now explains the guard and that it is not yet verified in two-player.
  • Minor (determinism): the earlier suggestion of a time-ranked mode that patches mid-game is withdrawn. It would move one player's RNG during play. CLAUDE.md forbids mid-game patching even with the guard.
  • Minor (cross-mod): compare with AutoWatch off versus on reported the watched patch methods as "Time B spends that A does not".
    • A method that only one side watched is now reported in a separate sentence: its time is inside other rows, and the Watch/AutoWatch settings differ.
    • Section 4 notes that [method] rows are nested.
    • The report's "no rows" line no longer blames Watch entries.
  • Minor (cross-mod): replacing prefixes. Such a prefix (LGP's PlantWater.TickPrefix, BeaverBuddies' TickBuckets prefix) is timed including the game's work that it does in the game's place.
    • A bool-returning prefix is now marked (can replace it) in its # watch| line and the report.
    • README and SESSION-README explain it, including that one on a tick loop holds nearly the whole tick.
  • Nit (cross-mod): the in-game Mod Settings note, the Settings.cs class doc, CLAUDE.md and the SessionService comment now list AutoWatch among the cfg-only, restart-needed settings. README "Working with other mods" says what the auto watch adds.
  • Minors (is the test real): the untested paths are now covered.
    • AutoInstall's free-slot count and Watch-entry names (M1, M2): EntriesKeepTheirSlots.
    • The Before tie-break (M8): OrderIsByName.
    • AutoFinal's never-called branch (M5): WatchedPatchIsTimed.
    • AutoLines and AutoCapability output (M7): WatchedPatchIsTimed.
    • summary.md still cut method names to Prefix(...): now Summary.ShortMethod plus a SummaryTests check.
    • Compare-mode naming (P1, P2): the new compare test.
  • Nits (is the test real):
    • The __originalMethod key check's message no longer claims more than .NET 8 can show.
    • A real TickableComponent subclass (Walker.Tick) is now checked (M4).
    • AutoFinal and the docs now say a never-seen method may be called only off the game thread.

Declined or left for a human

  • Nit (determinism): move AutoInstall to the Game-context configurator to close the background-thread window. Not done. The configurator runs before other mods' Load-time patches exist, so it would miss them, and the RNG restore already fixes the desync whatever the timing. The remaining risk is theoretical: another mod's own background thread entering a patch method while MonoMod writes its detour. It applies to any runtime Harmony patching, and it would show as a crash, not a silent desync. Recorded for a human.
  • "Consider leaving out the tick-loop methods" (Ticker.Update, TickBuckets, TickAll). Left as a decision. Those prefixes are now marked (can replace it) and documented. With this user's mods they fall outside the 40 slots anyway. Excluding them changes what the feature reports, so it is the owner's call.
  • M6 (Session.Start calls AutoInstall) and the Mono-only GetMethodFromHandle key (M3): these cannot be exercised in the .NET 8 test process, because Harmony cannot patch there. In-game item 9 covers them.

Open decisions for the owner

  • AutoWatch is off by default, a departure from 0.1.3's "most detail by default". Every call of a watched method pays for the wrapper, and some candidates run tens of thousands of times a second. PL2 (per-call patch cost not bounded by the budget) still waits on the A/B measurement.
  • The choice is by name at first load, not by measured time. Ranking by time would mean patching mid-game, which is now ruled out.
  • The 40-slot cap binds on the real setup. Session 2026-09-21_23-13-11 has about 78 to 81 candidates. The header lists every one left out, and a Watch entry can name any of them. The alternatives are a larger cap for the auto watch, or a hand-curated busiest-first tier.

🤖 Generated with Claude Code

kramsey458 and others added 2 commits September 22, 2026 08:06
…h slots

A mod's prefix on every entity's tick (BeaverBuddies) or on a singleton's
Tick (Late Game Performance) runs inside a profile row that names the game
or the singleton, so its cost never shows. Watch could time it, but only
if the user named each patch method.

New setting AutoWatch (PerformanceLog.cfg only, off by default). When the
first log of a run starts (every mod has started and the game has loaded,
so Harmony's registry holds their patches, and nothing ticks yet), the mod
reads the registry and puts its own Watch prefix and postfix around other
mods' prefixes, postfixes and finalizers on hot methods. It fills only the
Watch slots the Watch entries left (40 methods in all). It never unpatches,
reorders or changes another mod's patch, and never calls UnpatchAll.

The choice is pure code (AutoWatch.Plan) and goes in name order, whatever
order Harmony lists the patches in. It takes the patches on the methods
behind the profile's own rows first (a singleton's Tick, UpdateSingleton,
LateUpdateSingleton or StartParallelTick, or an entity's or component's
Tick), then the rest of PatchFormat's hot list. Within each group it goes
by patched method, then prefix, postfix, finalizer, then patch method. One
patch method is one watch, however many hot methods it is on. A method a
Watch entry already watches is not watched twice. A refused or failed
patch takes no slot. It does not rank by measured time: that would mean
patching while the game runs, which this mod avoids (parallel-tick work
spans frames).

The header gets # capability|autoWatch| and one # watch|...|auto|<kind on
target>|<owner> line per patch method, including the ones left out and
why. The end of the log gets # capability-final|autoWatch| with how many
were seen called. A tiny patch method that the runtime inlined into its
target cannot be seen, and a zero there is not a measurement. The frame
path is unchanged and allocation-free, and a check measures that for a
watched patch method.

perflog.py names a watched method by Class.Method(...) instead of the bare
"Prefix(...)", and says which hot method each auto-watched one is on.

Tests: AutoWatchTests (cap, Watch entries win, name order under any input
order, examining patches against real methods, and the wiring with no
allocation), config checks, and a test_perflog report check. C# 99 -> 104,
Python 46 -> 47. Fixtures' README.md and columns.md regenerated for the
method kind's new text. Docs: README, SESSION-README, CLAUDE.md,
TESTING.md (in-game check 9), packaging cfg.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Review found a co-op desync. The auto watch makes its Harmony patches at
the first PostLoad, after BeaverBuddies has seeded UnityEngine.Random for
the game (DeterminismService's constructor). Every Harmony.Patch builds a
MonoMod DynamicMethodDefinition, whose field initializer calls
Guid.NewGuid(), and BeaverBuddies' GuidPatcher turns each NewGuid into 16
draws from UnityEngine.Random, the state Timberborn's RandomNumberGenerator
uses. So a player with AutoWatch = true (or with a different number of
patches made) would start the first tick with a different random state.

Watch.AutoInstall now reads Harmony's registry and makes every patch
inside Watch.AroundAutoPatching, which in the game is KeepUnityRandom: it
reads UnityEngine.Random.state before anything runs and writes it back in
a finally, so the patches leave the game's random numbers exactly as they
found them. If the state cannot be read, nothing is patched. The config
Watch entries are patched at StartMod, before any seed, and are unchanged.
CLAUDE.md now says that any patch after a game starts loading must run in
that guard, and that nothing may be patched mid-game.

Review follow-ups in the same change:
- A prefix that returns bool (it can skip the method it patches) is marked
  "(can replace it)" in its # watch| line and the report, and README and
  SESSION-README say its time is work done instead of the game's.
- perflog.py compare no longer counts a method only one side watched as
  time that side spends (its time is inside the row of what runs it); it
  says the Watch or AutoWatch settings differ, and notes [method] rows in
  section 4. The report's "no rows" line no longer blames Watch entries
  when only the auto watch was on.
- summary.md names watched methods by Class.Method(...) (Summary.
  ShortMethod), as perflog.py already does.
- The in-game settings note, Settings.cs and CLAUDE.md list AutoWatch
  among the cfg-only settings; README "Working with other mods" says what
  the auto watch adds.
- AutoFinal and the docs say a never-seen method may be called only off
  the game thread.

Tests: GameRandomIsKept (the default guard is KeepUnityRandom, whose IL
reads the state before its try and writes it in its finally; every patch
and the registry read happen inside the guard), EntriesKeepTheirSlots
(Watch entries' slots and names through AutoInstall), and more checks in
OrderIsByName (a patch method on several hot methods is placed by the
first of them under 20 shuffles; the replace mark), PatchesAreExamined
(bool prefix, Walker.Tick as a real component), WatchedPatchIsTimed
(header lines, the never-called list), SummaryTests.ShortNames, and two
test_perflog cases (AutoWatch off versus on in compare; the report's line
without rows). C# 104 -> 106, Python 47 -> 49. Fixture README.md files
regenerated for the SESSION-README text.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@kramsey458
kramsey458 changed the base branch from claude/sampling-calibration to main September 22, 2026 19:50
kramsey458 and others added 2 commits September 22, 2026 12:50
…watch from throwing

- AutoWatch.Plan no longer takes patches on RandomNumberGenerator, Guid
  and DateTime by itself (status 'left out: ...'). In name order System.*
  sorts first, so BeaverBuddies' determinism patches, among the most-called
  methods in the game, took the free slots ahead of Ticker.Update and
  TickBuckets. A Watch entry can still name one.
- WatchPrefix/WatchPostfix now sit inside other mods' hot patch methods,
  so both are wrapped in try/catch and tolerate a null __originalMethod.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@kramsey458
kramsey458 marked this pull request as ready for review September 22, 2026 19:50
@kramsey458
kramsey458 merged commit 3e7c293 into main Sep 22, 2026
2 checks passed
@kramsey458 kramsey458 mentioned this pull request Sep 22, 2026
@kramsey458
kramsey458 deleted the claude/auto-watch branch September 22, 2026 19:54
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant