Skip to content

Latest commit

 

History

18 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Arabic-304M-Base

An Arabic-first decoder-only language model built from random weights in plain PyTorch. The project follows strict stage gates: implementation alone is not a pass; tests and measurable evidence are required before the next stage starts.

Current scope

Stages 0 and 1 are implemented and verified. Stage 2 engineering is implemented and tested, including conservative Arabic normalization, document provenance, real SentencePiece Unigram training, four-candidate bake-offs, held-out contamination checks, metrics, ranking, and safe freezing. Stage 2 now also has an audited source registry, hash-pinned Wikimedia acquisition, bounded XML and wikitext extraction, source-policy enforcement, and an 80-case balanced evaluation candidate with optional independent-review metadata. A strict read-only preflight now verifies the corpus, line-level provenance, hashes, evaluation integrity, and quarantine before any 16K/24K/32K/48K training begins. The production bakeoff completed, and the architecture-compatible 32K candidate is frozen with zero unknown tokens and exact Unicode round trips. Its manifest records explicit best-effort acceptance because the approximate MSA and dialect fertility targets were not met.

Stage 3 is complete on the target machine: 2,886,509 documents became 993,451,009 tokenizer-bound tokens across 100 audited shards. Stage 4 now has a minimal one-command path that trains a resumable 33,692,928-parameter pilot on 300,000,000 tokens, exports Safetensors, and opens a local Arabic desktop UI.

See state.json for exact evidence, HELP.md for the cumulative step-by-step operator guide, VM_SPECS.md for the measured build environment, BASH_HISTORY.txt for the sanitized project command history, docs/ARCHITECTURE.md for the model design, and docs/TOKENIZER.md for the tokenizer contract. Source acquisition and license controls are documented in docs/CORPUS_SOURCES.md.

The final target is a roughly 304M-parameter Base model. No pretrained model weights are used, and instruction tuning is deliberately outside the Base-model build.

Next: train the pilot and open the UI

From the prepared Windows workspace:

powershell -ExecutionPolicy Bypass -File .\finish_to_ui\run_all.ps1

Rerun the same command after an interruption. Reopen a completed model with:

powershell -ExecutionPolicy Bypass -File .\finish_to_ui\open_ui.ps1

Rebuild and test from scratch

python -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e '.[dev]'
python scripts/detect_hardware.py
python -m pytest -q
python -m pytest -q tests/test_tokenizer.py
python scripts/wikimedia_tokenizer_corpus.py validate --registry data/tokenizer/source_registry.json
python scripts/run_stage2.py --help
python scripts/run_stage3.py --help
python finish_to_ui/train_pilot.py --help
python finish_to_ui/demo_ui.py --help
python train.py --steps 160 --output checkpoints/stage1_tiny
python benchmark.py --config configs/tiny.json --batch-size 1 --sequence-length 64 --warmup 1 --steps 3

See docs/BUILD.md for verified commands and stage limits.

scripts/run_stage2.py performs all large Stage 2 work on the operator's machine: snapshot discovery, resumable verified downloads, extraction, corpus preparation, the exact four-size bake-off, and winner freezing. It writes step-specific logs and work/stage2/stage2_run_state.json, so a failure is reported precisely and a rerun resumes completed work.

Security baseline

  • Final weights use Safetensors.
  • Loading arbitrary executable checkpoints is forbidden.
  • trust_remote_code is forbidden.
  • Data and artifact provenance will be hash-addressed.
  • Tokenizer evaluation data is quarantined from its training corpus by normalized document hashes.
  • Production corpus records must match a hash-bound source-policy registry; Wikimedia JSONL must also match a complete extractor report. Free-form license labels are insufficient.
  • Remote dumps remain temporary until their official SHA-256 is verified.
  • Production tokenizer training stops before model creation unless the input preflight passes; freezing rechecks source files, candidate artifacts, and selected metrics instead of trusting a stored pass flag.
  • The released Base model has no shell, network, filesystem, credential, or tool access.

Verified Stage 1 result

On the available 9-vCPU environment, the 99,648-parameter test model reduced loss from 2.8183 to 0.0391 in 160 steps and greedily generated:

العلم نور والعمل يبني المستقبل

The checked-in configs/tiny.json architecture contains 22,910,592 parameters. The final target config is analytically confirmed at 304,140,288 parameters.

About

No description, website, or topics provided.

Resources

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages