Archive v0.8: catalog discovery and reconciliation - #23
Merged
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Goal
Promote v0.7 unmatched first-party model IDs from transient collection diagnostics into a durable, auditable Catalog Discovery / Reconciliation layer.
Discovery persistence
Migration
0010_catalog_discovery_reconciliation.sqladds:catalog_discovery_candidates: current per-provider remote-ID summary;catalog_discovery_observations: append-only evidence for each collection in which the ID appeared;catalog_reconciliation_events: append-only operator decisions.Candidate summaries retain first/last seen provenance, observation count, latest provider snapshot ID, current state, and optional resolved canonical Model.
Reconciliation states
Current candidate states:
discoveredmatchedignoredpromotion_readyAppend-only actions:
match_existingignoremark_promotion_readyreopenA match must resolve to a Model owned by the same Provider. This is enforced in both TypeScript and PostgreSQL.
Reclassification never rewrites prior decisions.
Drift boundary
Reconciliation is deliberately separate from identity ingestion.
match_existingrecords the catalog decision only. It does not create/replace an execution binding, insert an alias-resolution observation, or manufacture a v0.6 drift event.Execution identity can still change only through the source-backed v0.6 observer path.
Collector integration
Every parsed model list now:
observeFirstPartyIdentity;Repeated observations aggregate on
(provider, remote_model_id)while preserving every collection-level evidence row.Catalog adapter boundary
HTTP request and parsing behavior is extracted behind
CatalogSourceAdapter.Current adapters:
openai_modelsfor OpenAI-compatible model-list JSON;snapshot_onlyfor source capture without identity derivation.This separates scheduling/retrieval from parser-specific provider behavior and keeps the drift engine provider-neutral.
Operator workflow
New commands:
npm run list:catalog-discoveries -w @modelapse/catalog-adminnpm run reconcile:catalog-candidate -w @modelapse/catalog-adminThe write path remains CLI/database-operator scoped in v0.8; no public Web mutation endpoint is introduced.
Coverage
Integration coverage verifies:
See
docs/catalog-reconciliation.md.