From first principles. From zero. From Rust.
Sanskrit: beginning β a ground-up language model system in Rust.
aarambh-studio is a decoder-only language model implementation built with Rust and Candle. The repository covers the full engineering path: tokenization, model construction, training, inference, quantization, adapter tuning, alignment, evaluation, multimodal input, safety, and an OpenAI-compatible server.
The production source release is v3.0.0, with hybrid Gated DeltaNet, DeepSeek Sparse Attention, fine-grained MoE with shared experts, Multi-Token Prediction (MTP), on-policy distillation, native quantization-aware training, native video/document input, bounded long-horizon tool-use chains, persistent forgetting diagnostics, and Max thinking mode (16,384-token budget). v4.0.0-alpha.1 begins the v4 arc with Multi-Head Latent Attention (Phase 41) β a third attention kind that compresses the KV cache into a single low-rank latent per token.
Important
This is a source and engineering project. It does not publish crates to crates.io and does not ship pretrained checkpoints, adapters, GGUF files, or compiled binaries. You must train a model or provide compatible weights.
| Area | Capabilities |
|---|---|
| Model | RMSNorm, RoPE, GQA, SwiGLU, KV cache, tied embeddings, Tiny to Large configs |
| Efficient architecture | YaRN/NTK/linear RoPE scaling, Gated DeltaNet, learned block-sparse DSA, Multi-Head Latent Attention (MLA), fine-grained MoE, MTP |
| Training | BPE data pipeline, AdamW, cosine schedule, gradient accumulation/clipping, checkpoint resume, BF16 CUDA, single-node multi-GPU, on-policy distillation, native INT4/INT8 QAT |
| Fine-tuning | SFT, LoRA, QLoRA, DoRA, QDoRA, VLM adapters, GRPO, DPO, QDPO, tool-call tuning |
| Inference | Greedy/sampled decoding, streaming, thinking budgets, external or one-checkpoint MTP speculation, tool grammar, caller-executed chains |
| Model formats | SafeTensors, INT8, GPTQ/AWQ INT4, GGUF, Hugging Face conversion, quantized KV cache |
| Evaluation | Perplexity, MMLU-lite, HellaSwag, GSM8K, HumanEval-lite, preference, recall, multimodal/tool scorecards, capability forgetting curves, and MoE routing drift |
| Vision | Frozen CLIP-style encoder, image/video/document fusion, temporal and 2D layout encoding, multimodal DoRA/QDoRA tuning |
| Runtime | CPU SIMD, Rayon attention, optional custom CUDA PTX kernels, Axum 0.8.9 HTTP/SSE server |
| Guardrails | Prompt-injection checks, jailbreak checks, PII redaction, output scanning, streaming token safety, audit logs |
| Self-learning | Opt-in critique, replay, verifier rewards, deferred CPU updates, CUDA vision mode, and post-commit forgetting probes |
The implementation history and proof obligations for each feature live in the roadmaps. This README focuses on building and using the project.
- Rust 1.89 or newer
- Linux or another platform supported by Candle
- A C/C++ build toolchain for the bundled OpenH264 decoder
- Optional NVIDIA GPU and CUDA toolkit for
--features cuda nvccavailable at build time for custom CUDA PTX kernels- Python 3 only for dataset preparation scripts
CPU builds do not require CUDA. Tiny smoke configurations are designed for local development; Medium and Large training require suitable GPU memory.
git clone https://github.com/AarambhDevHub/aarambh-studio.git
cd aarambh-studio
cargo check --workspace --all-targets --locked
cargo test --workspace --locked
cargo build --release --locked -p aarambh-studio
target/release/aarambh-studio --helpRun a two-step CPU training smoke test using the checked-in Tiny Shakespeare fixture:
target/release/aarambh-studio train \
--config configs/tiny_shakespeare_smoke.tomlTrain the normal Tiny recipe:
target/release/aarambh-studio train \
--config configs/tiny_shakespeare.tomlTraining creates a tokenizer, model checkpoints, optimizer state, and
latest.json/best.json pointers under the configured checkpoint directory.
Smoke checkpoints validate the pipeline; two optimizer steps are not enough to
produce useful language quality.
aarambh-studio train Pretrain or continue a configured model
aarambh-studio infer Generate text or answer an image/video-grounded prompt
aarambh-studio agent Orchestrate bounded caller-executed tool-use chains
aarambh-studio eval Run evaluation tasks and compare scorecards
aarambh-studio quantise Calibrate and export INT8/INT4 GGUF checkpoints
aarambh-studio convert Convert SafeTensors, GGUF, or Hugging Face layouts
aarambh-studio finetune Run SFT, adapters, GRPO, DPO, VLM, or merge workflows
aarambh-studio distill Train/evaluate on-policy or offline teacher distillation
aarambh-studio selflearn Manage replay and persistent self-learning state
aarambh-studio serve Start the OpenAI-compatible HTTP/SSE server
Use aarambh-studio <command> --help for the complete option set.
See the phase-specific docs for full walkthroughs with smoke fixtures:
| Workflow | Guide |
|---|---|
| Train (CPU/GPU, MTP, MoE, distillation, QAT) | aarambh-studio train --help + configs/ TOML examples |
| Inference (text, thinking, image, video, document) | docs/aarambh-studio-complete-guide.md |
| Tool-use agent chains | docs/phase37_agent.md |
| Video understanding | docs/phase35_video.md |
| Document understanding | docs/phase36_document.md |
| OpenAI-compatible server | docs/inference-server.md |
| Evaluation & forgetting diagnostics | docs/phase38_forgetting.md |
| Multi-Head Latent Attention (MLA) | docs/phase41_mla.md |
| Quantization (GPTQ, QAT, GGUF) | docs/phase34_qat.md |
| MLA hybrid attention & KV-cache report | aarambh-studio eval --kv-cache-report + docs/phase41_mla.md |
| Fine-tuning (SFT, adapters, GRPO, DPO) | aarambh-studio finetune --help |
| Self-learning | SELF_LEARNING_V3.md |
# Minimal train-smoke β infer flow
cargo build --release --locked -p aarambh-studio
target/release/aarambh-studio train --config configs/tiny_shakespeare_smoke.toml
target/release/aarambh-studio infer --config configs/tiny_shakespeare.toml \
--model checkpoints/tiny_shakespeare_smoke/best/model.safetensors \
--tokenizer checkpoints/tiny_shakespeare_smoke/tokenizer.json \
--prompt "Hello" --max-tokens 16 --greedy| Scale | Parameters | Hidden | Layers | Heads | KV heads | FFN | Base context | RoPE theta |
|---|---|---|---|---|---|---|---|---|
| Tiny | 25M | 384 | 8 | 6 | 2 | 1,024 | 512 | 10,000 |
| Small | 117M | 768 | 12 | 12 | 4 | 2,688 | 1,024 | 10,000 |
| Medium | 360M | 1,024 | 24 | 16 | 8 | 3,392 | 2,048 | 500,000 |
| Large | 1.3B | 2,048 | 24 | 32 | 8 | 6,656 | 4,096 | 500,000 |
All standard scales use a 32,000-token vocabulary, RMSNorm epsilon 1e-5,
GQA, SwiGLU, and tied embeddings. Long-context and hybrid variants are selected
through TOML without changing the base scale definitions.
| Mode | Budget (tokens) | Default temperature | Default top-p | Use case |
|---|---|---|---|---|
none |
0 | 0.70 | 0.90 | Simple/comparative evals, no reasoning overhead |
low |
256 | 0.75 | 0.92 | Quick factual recall, short-answer tasks |
medium |
1,024 | 0.80 | 0.95 | Standard reasoning, multi-step math/code |
high |
4,096 | 0.80 | 0.95 | Complex multi-step proofs, long analysis |
max |
16,384 | 0.85 | 0.97 | Hard problems unsolved by High (Phase 39) |
The thinking budget is a ceiling on the number of content tokens emitted inside
the <think> block before the controller force-closes it. The effective budget
is clamped to min(mode.budget(), max_new_tokens - reserve), so Max never
exceeds the configured generation limit. Sampling defaults are applied only
when the caller does not supply explicit sampling parameters.
All five modes share the same ThinkingController mechanism β no structural
changes between modes. Every CLI command (infer, agent, serve, eval,
finetune grpo, distill train, selflearn start) accepts --thinking with
any of none, low, medium, high, or max. The server also accepts
reasoning_effort: "max" per request.
The workspace contains 18 internal library crates and one CLI package:
aarambh-studio-core Shared config, device, dtype, errors, and traits
aarambh-studio-tokenizer BPE tokenizer and reserved special tokens
aarambh-studio-data Datasets, preprocessing, sharding, and loaders
aarambh-studio-kernel CPU SIMD and optional CUDA kernels
aarambh-studio-nn Neural layers, attention, DeltaNet, DSA, MoE, and MTP
aarambh-studio-model Full decoder model and cache integration
aarambh-studio-weights SafeTensors, GGUF, conversion, and retrofit loading
aarambh-studio-quant INT8/INT4, GPTQ, AWQ, QAT, and KV quantization
aarambh-studio-train Optimizer, schedules, MTP loss, checkpoints, distributed train
aarambh-studio-finetune Adapters, SFT, GRPO, DPO, VLM, and tool tuning
aarambh-studio-inference Sampling, caching, thinking, MTP/external speculation, tools
aarambh-studio-agent Bounded tool chains, exact state, and caller-result ingestion
aarambh-studio-safety Input, output, streaming, PII, and audit policies
aarambh-studio-selflearn Critique, replay, verifiers, and persistent update state
aarambh-studio-eval Evaluation tasks, scorecards, and comparisons
aarambh-studio-vision Image/video/document decode, preprocessing, temporal/layout fusion
aarambh-studio-distill On-policy rollouts, teacher scoring, losses, and resume
aarambh-studio-serve Axum HTTP/SSE serving and continuous batching
aarambh-studio Command-line application
Packages inherit one workspace version and use publish = false.
Run the same primary gates used by CI:
cargo fmt --all --check
cargo check --workspace --all-targets --locked
cargo test --workspace --no-fail-fast --locked
cargo clippy --workspace --all-targets --locked -- \
-D warnings -D clippy::undocumented_unsafe_blocks
RUSTDOCFLAGS="-D warnings -D missing_docs" \
cargo doc --workspace --no-deps --locked
scripts/phase28_release_audit.shCUDA checks require a CUDA-capable environment and are intentionally opt-in.
| Document | Purpose |
|---|---|
| ARCHITECTURE.md | v1 model, training, inference, safety, and self-learning design |
| ARCHITECTURE_V2.md | v2 long context, vision, MoE, distributed, tools, and serving additions |
| ARCHITECTURE_V3.md | v3 hybrid attention, DSA, fine-grained MoE, MTP, agents, and forgetting diagnostics |
| ROADMAP.md | Completed v1 phases |
| ROADMAP_V2.md | Completed v2 phases through the v2.0.0 release |
| ROADMAP_V3.md | Current v3 delivery plan and status |
| SELF_LEARNING.md | Text self-learning design |
| SELF_LEARNING_V2.md | Vision-aware self-learning design |
| SELF_LEARNING_V3.md | v3 self-learning and forgetting-diagnostic integration |
| docs/aarambh-studio-config-toml-guide.md | Configuration field reference |
| docs/aarambh-studio-complete-guide.md | Beginner-oriented project walkthrough |
| docs/aarambh-studio-math-formulas-guide.md | Mathematical foundations and worked examples |
| docs/inference-server.md | Server endpoints, SDK usage, auth, safety, and limits |
| docs/phase32_mtp.md | MTP training, retrofit, exact speculation, and benchmark method |
| docs/phase33_distillation_results.md | On-policy distillation design, smoke proof, and comparison method |
| docs/phase34_qat.md | Native QAT configuration, continuation, export, and robustness validation |
| docs/phase35_video.md | Video migration, decoding, tuning, inference, and NExT-QA evaluation |
| docs/phase36_document.md | PDF/page ingestion, layout tuning, inference, and DocVQA ANLS evaluation |
| docs/phase37_agent.md | Tool-chain protocol, safety, context policy, SFT, and response-path evaluation |
| docs/phase38_forgetting.md | Capability curves, routing drift, training/self-learning hooks, and Manas JSONL |
| RELEASE.md | Source-release process and artifact policy |
| CHANGELOG.md | Versioned implementation history |
- No pretrained model, GGUF, adapter, or binary ships β you train your own.
- MoE uses dense masked dispatch (not sparse grouped). Multi-GPU is single-node.
- Tool chains are generated and orchestrated but never executed by the runtime.
- Video is visual-only H.264 MP4; audio is unsupported.
- Documents are pixel-based (no OCR/table parser).
- The server is local/single-model; vision and self-learning are CLI workflows.
Full exclusions in the versioned roadmaps.
Read CONTRIBUTING.md before opening a pull request. Use GitHub issues for reproducible bugs and scoped feature requests. Report vulnerabilities through SECURITY.md, not a public issue.
@software{aarambh_ai_2026,
title = {aarambh-studio: A Ground-Up Language Model System in Rust},
author = {Aarambh Dev Hub},
year = {2026},
url = {https://github.com/AarambhDevHub/aarambh-studio},
version = {3.0.0},
license = {Apache-2.0}
}Licensed under the Apache License 2.0.