- Checkpoints use Safetensors, never pickle-backed
torch.load. - Configuration and metadata use strict JSON.
- Unknown configuration fields and invalid head relationships fail early.
- A model-only checkpoint has exactly
model.safetensors,config.json, andmanifest.json; unexpected manifest paths are rejected. - File sizes and SHA-256 hashes are verified before tensors load.
- Declared model size is checked before weight allocation; the default loading ceiling is 500M parameters.
- Tied embedding storage is represented once in the weight file.
- The model core has no shell, network, filesystem, credentials, or tools.
- Tokenizer input is strict UTF-8 JSONL with rejected duplicate keys, unknown fields, invalid categories, missing source IDs, and missing licenses.
- Every normalized tokenizer document has a SHA-256 and line-level provenance.
- Exact normalized evaluation/training overlap blocks the bake-off.
- SentencePiece models are capped by byte size and inspected as protobuf before loading; Unigram, byte fallback, identity normalization, and special IDs are mandatory.
- A read-only production preflight verifies original input and extractor-report hashes, source policy, corpus/provenance hashes, every line-to-hash binding, counts, source balance, evaluation integrity, and quarantine before any four-size training starts.
- Freezing never trusts only a stored pass flag: it reruns preflight, verifies all candidate artifacts, and recomputes every candidate's metrics and ranking.
- A strict source registry binds IDs, licenses, categories, attribution rules, and allow/block decisions into the prepared-corpus manifest.
- Wikimedia acquisition permits only HTTPS under the official current-content
dump tree, rejects cross-host redirects, requires an identifying user agent,
and pins official
SHA256SUMSbefore download. - Downloads are size bounded, atomic, sequential, and hash verified before hardened XML parsing or wikitext processing.
- Prepared Wikimedia JSONL is accepted only when a complete extractor report binds its source identity, registry, importer version, counts, and SHA-256.
- Production evaluation requires balanced scale, unique IDs/text, a matching manifest, and quarantine from training data. Human linguistic review remains useful advisory evidence.
- Stage 3 rejects invalid UTF-8 and duplicate JSON keys, high-confidence credentials, spam anomalies, exact duplicates, and near duplicates.
- Direct identifiers are deterministically redacted; every clean record keeps source, license, document identity, content hash, and original provenance.
- Domain balance, input hashes, normalizer/filter versions, tokenizer identity, shard sizes, token/document counts, and SHA-256 are manifest-bound.
- Token shards use non-executable uint32/uint64 binary files and verified, safe-basename metadata; memory-mapped reads occur only after validation.
SHA-256 proves that bytes match the manifest; it does not prove who produced the manifest. Artifact signing and publisher identity remain release-stage controls. Do not load checkpoints from an untrusted publisher merely because their internal hashes are consistent.
Do not add:
torch.load(untrusted_path)
pickle.load(...)
trust_remote_code=True
dynamic import from checkpoint metadata
shell commands from configuration
Full benchmark quarantine, dependency lock hashes, signing, and training isolation remain Stage 6/7 work. PII/secret filtering, anomaly checks, dedup, domain-cap accounting, immutable clean data, and shard verification are now implemented and fixture-tested in Stage 3.