Skip to content

Archive v0.8: catalog discovery and reconciliation - #23

Merged
WangEn merged 20 commits into
mainfrom
archive/v0.8-catalog-reconciliation
Oct 1, 2026
Merged

WangEn merged 20 commits into
mainfrom
archive/v0.8-catalog-reconciliation

Conversation

@WangEn

@WangEn WangEn commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Goal

Promote v0.7 unmatched first-party model IDs from transient collection diagnostics into a durable, auditable Catalog Discovery / Reconciliation layer.

Discovery persistence

Migration 0010_catalog_discovery_reconciliation.sql adds:

  • catalog_discovery_candidates: current per-provider remote-ID summary;
  • catalog_discovery_observations: append-only evidence for each collection in which the ID appeared;
  • catalog_reconciliation_events: append-only operator decisions.

Candidate summaries retain first/last seen provenance, observation count, latest provider snapshot ID, current state, and optional resolved canonical Model.

Reconciliation states

Current candidate states:

  • discovered
  • matched
  • ignored
  • promotion_ready

Append-only actions:

  • match_existing
  • ignore
  • mark_promotion_ready
  • reopen

A match must resolve to a Model owned by the same Provider. This is enforced in both TypeScript and PostgreSQL.

Reclassification never rewrites prior decisions.

Drift boundary

Reconciliation is deliberately separate from identity ingestion.

match_existing records the catalog decision only. It does not create/replace an execution binding, insert an alias-resolution observation, or manufacture a v0.6 drift event.

Execution identity can still change only through the source-backed v0.6 observer path.

Collector integration

Every parsed model list now:

  1. sends exact current first-party binding matches through observeFirstPartyIdentity;
  2. persists all remaining IDs as durable Discovery Candidates + append-only observations;
  3. keeps bounded candidate IDs and counts in collection metadata for diagnostics.

Repeated observations aggregate on (provider, remote_model_id) while preserving every collection-level evidence row.

Catalog adapter boundary

HTTP request and parsing behavior is extracted behind CatalogSourceAdapter.

Current adapters:

  • openai_models for OpenAI-compatible model-list JSON;
  • snapshot_only for source capture without identity derivation.

This separates scheduling/retrieval from parser-specific provider behavior and keeps the drift engine provider-neutral.

Operator workflow

New commands:

  • npm run list:catalog-discoveries -w @modelapse/catalog-admin
  • npm run reconcile:catalog-candidate -w @modelapse/catalog-admin

The write path remains CLI/database-operator scoped in v0.8; no public Web mutation endpoint is introduced.

Coverage

Integration coverage verifies:

  • the same unknown remote ID aggregates across two first-party collections;
  • collection snapshots still feed known binding identity drift;
  • match / ignore / promotion-ready / reopen transitions append audit events;
  • cross-provider matches fail;
  • direct SQL cross-provider reconciliation fails at the database trigger;
  • discovery observations and reconciliation decisions are append-only.

See docs/catalog-reconciliation.md.

@WangEn
WangEn merged commit bbabb14 into main Oct 1, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant