TinyTitan runs Qwen-family MoE and dense text models locally on Apple Silicon, streaming the routed experts from SSD so a model larger than your RAM still runs. Decode is native Swift over hand-written Metal kernels; the prompt runs on the GPU or, optionally, the Apple Neural Engine. It ships a CLI, an installer/repacker, and a loopback OpenAI-compatible server with no MLX and no GGUF dependency.
Who it is for. People on an M-series Mac who want to run a 35B — or a 125B —
model locally on a machine that could never hold it in RAM, and developers who
want a loopback OpenAI- or Anthropic-compatible endpoint for Codex, Claude Code,
Qwen Code, OpenCode, Zed or the official SDKs. It is a terminal-first tool: the
engine and its server are the product, and a browser chat window is an optional
client of that server. It is not a fine-tuning toolkit, a vision model, a
GUI app, or a way to expose a model to your network — the server is
127.0.0.1-only and has no authentication.
- The LLM engine as a library —
TinyTitanLib. Embed the engine in your own Swift program: anEngineand aSession, streaming tokens through a callback. Same kernels, format reader and sampler as the engine, in your process — no subprocess, no HTTP, no model reimplementation. - The complete local LLM engine —
TinyTitanCLI+TinyTitanServer. Run a 35B or 125B model from the terminal and serve it on a loopback OpenAI-/Anthropic-compatible endpoint for Codex, Claude Code, Qwen Code, OpenCode, Zed or the official SDKs. Native Swift and Metal: no MLX, no GGUF. - DeepSeek Harness plugins —
plugins/dsh-tinytitan,plugins/dsh-lan-manager. The harness finds the models you have installed and keeps its route current, with a quiet compaction backend; the LAN manager drives a fleet of harnesses from one console. - A ready-made DSH bundle install —
tools/dsh_local.sh. A pinned harness with the TinyTitan bundle already wired, entirely under~/.tinytitan: it never touches adsh, a~/.dshor a port you already use. - The
.ssdaimodel format and its converter —TinyTitanRepack+tools/. Stream a checkpoint into a hash-verified install that keeps the routed experts on SSD, so a model larger than your RAM still runs — and a truncated, swapped or edited install is refused rather than served.
Those five ship from one branch (main) and one release tag. The first two are
the products — a library and an engine — and the rest are the bundles and
tooling that ship beside them; the split is spelled out under
Library and engine.
Peak decode on a base 8-core M3 MacBook Pro with 24 GB. NA means the CPU
engine does not serve that model: the MoE families stream their experts on the
GPU + ANE path, and only the dense Qwen 3.5 models run on either engine.
Per-version measurements and the method live on the wiki
(Benchmarks ·
Benchmarking Guide).
| Model | Quantization | GPU | CPU |
|---|---|---|---|
| Qwen 3.5 2B (dense) | 4-bit | 53.73 tok/s | 15.42 tok/s |
| Qwen 3.5 2B (dense) | 8-bit | 32.77 tok/s | 15.83 tok/s |
| Qwen 3.5 4B (dense) | 4-bit | 26.18 tok/s | 7.71 tok/s |
| Qwen-AgentWorld 35B-A3B | 4-bit | 21.74 tok/s | NA |
| Ornith 1.5 35B-A3B | 4-bit | 21.65 tok/s | NA |
| Qwen 3.6 35B-A3B | 4-bit | 21.41 tok/s | NA |
| KAT-Coder-V2.5-Dev 35B-A3B | 4-bit | 17.86 tok/s | NA |
| Qwen 3.5 4B (dense) | 8-bit | 16.14 tok/s | 7.04 tok/s |
| Qwen 3.5 9B (dense) | 4-bit | 14.93 tok/s | 4.07 tok/s |
| Qwen 3.6 35B-A3B | 8-bit | 12.37 tok/s | NA |
| Qwen-AgentWorld 35B-A3B | 8-bit | 12.28 tok/s | NA |
| Ornith 1.5 35B-A3B | 8-bit | 11.93 tok/s | NA |
| Qwen 3.5 9B (dense) | 8-bit | 8.90 tok/s | 4.51 tok/s |
| KAT-Coder-V2.5-Dev 35B-A3B | 8-bit | 6.91 tok/s | NA |
| Qwen3.8-Flash-Next 125B-A6B | 4-bit | 5.46 tok/s | NA |
| Qwen3.8-Flash-Next 125B-A6B | 8-bit | 2.10 tok/s | NA |
- Getting started
- Installation and configuration — the model catalogue, environment variables, the launcher and the tools
- Cookbook — one recipe per task
- Local server and API
- Runtime controls — every flag and default
- Features
- Library and engine — the library and the engine, and how the two products are split
- System design
- FAQ
- Benchmarks · Benchmarking guide
- Changelog
- Project tracker — open work only
- Repository layout — where everything lives, and the naming and file-size conventions
TinyTitan is a focused fork of drumih/turbo-fieldfare, whose bounded-memory runtime, installer, CLI and local server this project builds on. The Qwen 3.6 integration was created by NeelM0906 in upstream PR #29. Concise mode is derived from the Nail-Qwen3.6-35B-A3B chat template by peculiar-ragdoll.
Related research: TinyTitan Datacenter runs large MoE models across a cluster of Mac minis and Studios, keeping them on SSD/NVMe for near-linear scale of decode throughput.
Apache License 2.0 — see LICENSE and NOTICE. Copyright (c) 2026 André Borchert.
