Skip to content

A15-05: the custom-TreeScorer API is an untested facade — Ratchet() never forwards TreeScorer to its Bootstrapper, so the two search phases optimise different objectives #137

Description

@ms609

Severity: sev:high · Area: 15 (Legacy pure-R search API)

Covers red-team candidates A15-O-02 (high), A15-O-03 (med), A15-O-11 (med).

The custom-criterion API is this area's stated reason to exist — TreeSearch()/Ratchet() take a TreeScorer argument precisely so a user can optimise something other than equal-weights parsimony. Three independent defects show that path is not exercised end to end.

1. Ratchet() never passes TreeScorer to its Bootstrapper — the two phases optimise different objectives

R/Ratchet.R:145-150 calls

Bootstrapper(edgeList, initializedData, maxIter = , maxHits = , verbosity = ,
             EdgeSwapper = , stopAtPeak = , stopAtPlateau = , ...)

with no TreeScorer =. R/Bootstrap.R:27-30 then calls EdgeListSearch() without one, so the perturbation phase falls back to the default EdgeListScore.

The decisive point is that ... cannot rescue this. TreeScorer is a named formal of Ratchet() (R/Ratchet.R:82), so argument matching consumes it and it can never reach Bootstrapper through ....

Confirmed behaviourally by instrumenting the Bootstrapper: the named arguments arriving were exactly maxIter, maxHits, verbosity, EdgeSwapper, stopAtPeak, stopAtPlateau, and the user's custom scorer was called 0 times during each bootstrap phase against 391 calls overall.

Jackknife() forwards it correctly (R/Jackknife.R:81), which is what makes this an oversight rather than a design.

Consequence: the ratchet perturbs under equal-weights parsimony and then searches under the user's criterion. The whole point of the ratchet is that the perturbed landscape guides the search; perturbing under a different objective silently degrades it to something closer to random restarts.

2. Ratchet() with a custom TreeScorer and a non-ParsimonyData dataset cannot run at all

R/Bootstrap.R:19-21 does startWeights <- dataset[["original_weight"]], which is NULL for a raw phyDat, then rep.int(eachChar, startWeights):

Error: invalid type (NULL) for 'times' (must be a vector)

The documented contract really is that broad — R/Ratchet.R:12-13:

a dataset in the format required by TreeScorer(); the default EdgeListScore() requires a ParsimonyData object

so a user supplying a scorer that wants a plain phyDat is following the documentation exactly, and cannot run. Note the internal contradiction: R/Bootstrap.R:3 documents the same argument as "A ParsimonyData object". One of the two is wrong.

Severity note: the finder proposed high; the verifier downgraded this to med because it is an immediate hard error, not a silent wrong answer. Recorded as med here.

3. SuccessiveWeights() errors on every input it accepts

R/SuccessiveApproximations.R:175-189 converts phyDatsaDat via PrepareDataSA(), then calls CharacterLength(tree, dataset, compress = TRUE). CharacterLength() opens with .CheckDataCharLen(), which hard-requires phyDat:

Error: Dataset must be of class phyDat, not saDat.

The function is exported and documented (@rdname SuccessiveApproximations) as returning a tree score. It has no test coverage, which is why total breakage went unnoticed.

This also answers a question the scope row has carried since this area was created — whether SuccessiveWeights() has drifted from the C++ IW kernel. It cannot be compared, because it never returns a value.

The pattern, which is the real finding

Taken with #126 (SuccessiveApproximations() lacks the sectorial/fuse/pool machinery, undocumented) and #131 (eight contract drifts), the picture across this area is consistent: the advertised custom-criterion API has decayed into an untested façade over the native equal-weights kernel. Each defect individually looks minor; together they mean no user has successfully driven a custom scorer through Ratchet() in a long time.

The highest-value follow-up is therefore not more per-function bug hunting but a deliberate audit of which advertised custom-criterion entry points actually work end to end, with a test for each that uses a scorer distinguishable from EdgeListScore — the defect in §1 is invisible to any test whose custom scorer happens to agree with the default.

Verification

§1 and §2 confirmed REAL by an independent peer-tier verifier against commit 826d332b0, §1 both statically ("TreeScorer" %in% names(formals(Ratchet)) is TRUE) and behaviourally (call-count instrumentation). §3 confirmed by a separate verifier reproducing the error.

Found by /red-team area 15, 2026-08-05, opus (Opus 5). Not found by a sonnet pass over the same files earlier the same day.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:15Red-team focus area 15red-teamFiled by the /red-team rotationsev:highP1: wrong user-visible result / crash

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions