Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
20 commits
Select commit Hold shift + click to select a range
7fbe1f7
Archive v0.8: add catalog discovery and reconciliation persistence
WangEn Oct 1, 2026
8d58281
Archive v0.8: implement discovery candidate reconciliation
WangEn Oct 1, 2026
0cd64ec
Archive v0.8: add catalog source adapter interface
WangEn Oct 1, 2026
f623f47
Archive v0.8: share discovery model type through adapter
WangEn Oct 1, 2026
80f2965
Archive v0.8: persist discoveries through catalog adapters
WangEn Oct 1, 2026
df3bf1a
Archive v0.8: export catalog discovery and adapters
WangEn Oct 1, 2026
8cff138
Archive v0.8: add discovery reconciliation operator CLI
WangEn Oct 1, 2026
e8570e2
Archive v0.8: expose discovery operator scripts
WangEn Oct 1, 2026
88c95fd
Archive v0.8: enforce reconciliation provider invariants
WangEn Oct 1, 2026
27a99c2
Archive v0.8: harden discovery chronology and actions
WangEn Oct 1, 2026
f51c880
Archive v0.8: expose parser through adapter module
WangEn Oct 1, 2026
a4fc1a3
Archive v0.8: cover catalog adapter contract
WangEn Oct 1, 2026
7ca4a88
Archive v0.8: verify discovery aggregation and reconciliation audit
WangEn Oct 1, 2026
5bd8581
Archive v0.8: document catalog discovery reconciliation
WangEn Oct 1, 2026
1e22ce5
Archive v0.8: index current reconciliation by insertion order
WangEn Oct 1, 2026
b97871d
Archive v0.8: project latest reconciliation by audit insertion
WangEn Oct 1, 2026
118eca4
Archive v0.8: bound discovery metadata projection
WangEn Oct 1, 2026
2e60d1d
Archive v0.8: avoid duplicate adapter model export
WangEn Oct 1, 2026
33ee885
Archive v0.8: validate discovery filters and ids
WangEn Oct 1, 2026
a701245
Archive v0.8: type reconciliation timestamp assignment
WangEn Oct 1, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
173 changes: 173 additions & 0 deletions docs/catalog-reconciliation.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,173 @@
# Catalog Discovery / Reconciliation

Archive v0.8 promotes unmatched first-party model IDs from collection-run diagnostics into an auditable discovery queue.

The discovery layer does **not** auto-create Models, rewrite execution bindings, or synthesize identity drift. It separates three facts:

1. a first-party source exposed a remote model ID;
2. Modelapse has or has not reconciled that ID to a canonical Model;
3. a source-backed execution identity observation changed over time.

Only (3) feeds the v0.6 drift engine.

## Data model

### `catalog_discovery_candidates`

One mutable summary row per `(provider, remote_model_id)`.

It keeps:

- first / last seen timestamps;
- first / last collection and source references;
- latest explicit provider snapshot ID when supplied by the source;
- observation count;
- current reconciliation state;
- optional resolved canonical Model.

Current states:

```text
discovered
matched
ignored
promotion_ready
```

The summary is an index over immutable history; it is not the audit log.

### `catalog_discovery_observations`

Append-only evidence that a candidate appeared in a specific first-party collection run.

Repeated model-list collections therefore answer both:

- “is this ID still present?”
- “how many independent retrievals have observed it?”

### `catalog_reconciliation_events`

Append-only operator decisions:

```text
match_existing
ignore
mark_promotion_ready
reopen
```

A `match_existing` decision must resolve to a Model owned by the same Provider. This is enforced both by the TypeScript repository and a PostgreSQL trigger.

Reclassification writes another event; prior decisions are never rewritten.

## Collector integration

For every parsed model list, v0.8 divides remote IDs into two sets:

```text
exact current first-party binding match
-> v0.6 observeFirstPartyIdentity(...)
-> alias / binding history and drift

not an exact current binding match
-> Catalog Discovery Candidate
-> append-only Discovery Observation
```

An unmatched ID can remain visible for many collections without becoming canonical identity.

Collection-run metadata still contains `unmatchedRemoteModelIds` for lightweight diagnostics, and now also records `discoveryCandidateIds` so the run can be joined directly to the durable discovery queue.

## Reconciliation semantics

### Match existing

Use when evidence supports that the remote ID belongs to an already-cataloged Model.

This records the reconciliation only. It does **not**:

- add or replace an execution binding;
- insert an alias-resolution event;
- close a previous binding;
- generate a v0.6 drift event.

Those operations require a source-backed identity observation with the normal v0.6 chronology and immutability rules.

### Ignore

Use for IDs that should remain observed evidence but are not useful as a Modelapse canonical Model candidate, for example non-chat utility endpoints or provider-internal artifacts.

Future collections continue to update first/last-seen evidence without erasing the ignore decision.

### Promotion ready

Marks a discovery as ready for a later explicit canonical registration workflow.

v0.8 deliberately stops before automatic promotion. Canonical Model creation still requires a deliberate registration step with sourced identity fields.

### Reopen

Returns any reconciled candidate to `discovered` while preserving every earlier decision event.

## Catalog adapter boundary

The HTTP / parsing behavior is now behind `CatalogSourceAdapter`.

Current adapters:

- `openai_models`: OpenAI-compatible JSON model-list parsing;
- `snapshot_only`: capture and hash the source without deriving model identities.

The Observer selects adapters by the source row's `parser` value. Provider/source scheduling remains data-driven in `catalog_observer_sources`.

This keeps scheduling, retrieval, parsing, discovery, reconciliation, and identity ingestion as separate layers so provider-specific parsing can expand without coupling to the drift engine.

## Operator CLI

List candidates:

```bash
npm run list:catalog-discoveries -w @modelapse/catalog-admin
```

Optional filters:

```text
MODELAPSE_PROVIDER_SLUG=deepseek
MODELAPSE_CATALOG_DISCOVERY_STATUS=discovered
```

Reconcile one candidate:

```bash
npm run reconcile:catalog-candidate -w @modelapse/catalog-admin
```

Required:

```text
DATABASE_URL=...
MODELAPSE_CATALOG_CANDIDATE_ID=<uuid>
MODELAPSE_RECONCILIATION_ACTION=match_existing|ignore|mark_promotion_ready|reopen
MODELAPSE_RECONCILIATION_ACTOR=<operator identity>
```

For `match_existing`:

```text
MODELAPSE_RESOLVED_MODEL_ID=<model uuid>
```

Optional:

```text
MODELAPSE_RECONCILIATION_NOTE=...
```

Operator identity is stored as audit data. Do not put secrets in the actor or note fields.

## Promotion remains explicit

A `promotion_ready` candidate is evidence for the next workflow, not a Model.

A later promotion layer should require an explicit canonical slug, marketing name, status/family/track decisions, source selection, and first-party execution binding before the candidate can become a canonical Model.
4 changes: 3 additions & 1 deletion packages/catalog-admin/package.json
Original file line number Diff line number Diff line change
Expand Up @@ -16,7 +16,9 @@
"bootstrap:deepseek-smoke": "node dist/src/cli.js bootstrap-deepseek-smoke",
"bootstrap:deepseek-flash-model": "node dist/src/cli.js bootstrap-deepseek-flash-model",
"observe:first-party-identity": "node dist/src/cli.js observe-first-party-identity",
"collect:first-party-catalog": "node dist/src/cli.js collect-first-party-catalog"
"collect:first-party-catalog": "node dist/src/cli.js collect-first-party-catalog",
"list:catalog-discoveries": "node dist/src/cli.js list-catalog-discoveries",
"reconcile:catalog-candidate": "node dist/src/cli.js reconcile-catalog-candidate"
},
"dependencies": {
"@modelapse/blob-store": "*",
Expand Down
106 changes: 106 additions & 0 deletions packages/catalog-admin/src/catalog-adapter.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,106 @@
export type CatalogObserverSourceKind = "model_list" | "docs";
export type CatalogObserverParser = "openai_models" | "snapshot_only";

export interface ObservedRemoteModel {
readonly id: string;
readonly providerSnapshotId: string | null;
}

export interface CatalogSourceAdapter {
readonly parser: CatalogObserverParser;
requestHeaders(input: {
readonly sourceKind: CatalogObserverSourceKind;
readonly credential: string | undefined;
readonly collectorBuild: string;
}): Readonly<Record<string, string>>;
parseModelList(body: string): readonly ObservedRemoteModel[] | null;
}

function isRecord(value: unknown): value is Record<string, unknown> {
return typeof value === "object" && value !== null && !Array.isArray(value);
}

function optionalSnapshotId(value: Record<string, unknown>): string | null {
for (const key of [
"provider_snapshot_id",
"providerSnapshotId",
"snapshot",
"version",
"model_version",
]) {
const raw = value[key];
if (typeof raw === "string" && raw.trim()) return raw.trim();
}
return null;
}

export function parseOpenAICompatibleModelList(
input: unknown,
): readonly ObservedRemoteModel[] {
if (!isRecord(input) || !Array.isArray(input.data)) {
throw new Error("Model-list payload must contain a data array");
}

const models = new Map<string, ObservedRemoteModel>();
for (const item of input.data) {
if (!isRecord(item) || typeof item.id !== "string" || !item.id.trim()) {
continue;
}
const id = item.id.trim();
models.set(id, {
id,
providerSnapshotId: optionalSnapshotId(item),
});
}
return [...models.values()].sort((a, b) => a.id.localeCompare(b.id));
}

function defaultHeaders(input: {
readonly sourceKind: CatalogObserverSourceKind;
readonly credential: string | undefined;
readonly collectorBuild: string;
}): Record<string, string> {
const headers: Record<string, string> = {
accept:
input.sourceKind === "model_list"
? "application/json"
: "text/html,application/xhtml+xml,text/plain;q=0.8,*/*;q=0.5",
"user-agent":
"modelapse-catalog-observer/" + input.collectorBuild.slice(0, 64),
};
if (input.credential) {
headers.authorization = "Bearer " + input.credential;
}
return headers;
}

const OPENAI_MODELS_ADAPTER: CatalogSourceAdapter = {
parser: "openai_models",
requestHeaders: defaultHeaders,
parseModelList(body) {
return parseOpenAICompatibleModelList(JSON.parse(body));
},
};

const SNAPSHOT_ONLY_ADAPTER: CatalogSourceAdapter = {
parser: "snapshot_only",
requestHeaders: defaultHeaders,
parseModelList() {
return null;
},
};

const ADAPTERS: Readonly<Record<CatalogObserverParser, CatalogSourceAdapter>> = {
openai_models: OPENAI_MODELS_ADAPTER,
snapshot_only: SNAPSHOT_ONLY_ADAPTER,
};

export function getCatalogSourceAdapter(
parser: CatalogObserverParser,
): CatalogSourceAdapter {
const adapter = ADAPTERS[parser];
if (!adapter) {
throw new Error("Unsupported catalog source parser: " + parser);
}
return adapter;
}
Loading
Loading