Skip to content

Experiment: Laya as a relevance check in front of recall #87

Description

@ashu17706

Question

Can a local System One model (Laya) replace retrieval, or only reject hits that are off-topic?

Jev signups are paused, so this run used the open-weight typed-decisions checkpoint instead. Jev itself was not called.

Setup

  • Model: aac6fef/laya-typed-decisions-mlx via laya-mlx 0.2.0, Apple Silicon, MLX 0.32.2.
  • Fixtures: the five CI recall probes in test/eval/fixtures/ (auth migration, deploy pipeline, density tie).
  • Baseline: recall() as the probes already run in test/recall-quality.test.ts.
  • Laya arm: show it every session in the scenario, not only the search hit. One question: "This session answers the query." Keep the session when P(true) ≥ 0.5. The threshold was fixed before looking at the scores.
  • The stored text is the original session. Laya does not rewrite it.
  • Score with the same recall / precision / substring checks as test/eval/fixtures/score.ts.

Result

Search passed 5/5. Laya passed 3/5. Warm calls were about 25 ms.

Probe Search Laya
JWT session cookies keeps the migration, drops the password note same (0.61 keep, 0.26 drop)
Mobile cookie migration keeps the follow-up keeps the follow-up and the original decision
Password hashing keeps only the bcrypt note same; confidence on the two drops was 0.82 and 0.86
Flaky CI pipeline project filter keeps the API service and drops the frontend keeps both (0.59 and 0.52)
Message queue density score keeps the high-signal copy both copies score 0.64, so it keeps both

The deploy miss was 0.52, just over the line. The line was not moved. The density pair would fail at any threshold, because the two texts are identical.

Reading

Laya is a useful reject step when the extra row is about something else. It is not a project filter, and it cannot break a tie between two sessions with the same text. Those two probes are why search scores 5/5 and Laya scores 3/5.

This does not say decision models fail at memory. It says this local model, on these five probes, did not beat the retrieval we already have.

Not in this run

An earlier check asked Laya to label 11 hand-labeled note pairs as supersedes, contradicts, relatesTo, or none. It got 6/11, with choice confidence from 0.01 to 0.36. That job was not repeated here.

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions