Train task-specific models on a connectome. Run them in your own application.
CNSKit connects MaleCNS graph data to a practical Python workflow: bind observations to selected neurons, train a readout or input adapter, evaluate held-out episodes, and export a stateful inference model. The graph stays sparse and its body IDs and source provenance travel with the model.
Quickstart | Training guide | MaleCNS data | API | Evidence
| Workflow | Implementation |
|---|---|
| Custom sequence regression or classification | Episode datasets, partial labels, early stopping, Adam and truncated backpropagation |
| Fixed-graph experimentation | Train only a readout, input/output adapters, or adapters plus bounded per-neuron gains and leak |
| MaleCNS integration | Verified graph import, exact body-ID binding, explicit induced subgraphs |
| Stateful inference | Batched sequences or one step at a time; state is passed explicitly |
| Reproducible export | NumPy checkpoints without pickle, graph fingerprint, ordered channels, fitted normalization and training report |
| Lightweight baseline | NumPy ridge readout fitting remains available without PyTorch |
The framework is task-general; it is not a pretrained general intelligence model. Anatomical connectivity is an experimental prior. Dynamics, observation encodings and learned outputs are model choices. We do not claim biological fidelity, a MaleCNS advantage over ordinary networks, or whole-connectome training speedups.
Python 3.11 or newer. No account, server or connectome download is needed for the synthetic example.
git clone https://github.com/OpenUploading/cnskit.git
cd cnskit
python -m venv .venv
# Linux/macOS: source .venv/bin/activate
# Windows: .venv\Scripts\Activate.ps1
python -m pip install -e ".[train,malecns]"
python examples/train_sequence.py --out runs/first-policyThis fits a temporal filtering task on a synthetic graph, evaluates separate test episodes, compares against observation-only ridge regression, and verifies checkpoint reload. It demonstrates the workflow, not MaleCNS performance. PyTorch is optional; install the appropriate CPU/CUDA wheel for your environment before the package if needed.
For masked sequence classification, run python examples/train_classification.py --out runs/classifier. For matched recurrent and zero-edge baselines, run python examples/compare_baselines.py.
from cnskit import load_policy
import numpy as np
policy = load_policy("runs/first-policy")
prediction, state = policy.predict(np.zeros((1, 16, 2), dtype="float32"))
# Continue the SAME episode by passing state; omit it to start a new episode.
prediction, state = policy.predict(np.ones((1, 8, 2), dtype="float32"), state=state)Prepare the official data using the dataset guide, then run:
python examples/malecns_task.py --graph /data/traced-graph --out runs/malecns-taskThis is real MaleCNS anatomy with a synthetic regression task. The script selects 128 neurons deterministically from the loaded graph, fits input/output adapters, and evaluates three initialization seeds against a zero-edge ablation using the same observations, splits and training budget. It exports the first requested seed, never the seed with the best test score.
| Output | What it contains |
|---|---|
task.json |
Actual selected body IDs and a reusable CLI configuration |
train.npz, validation.npz, test.npz |
Separate episodes with inputs, targets and IDs |
policy/ |
Graph-bound model, fitted normalization and training report |
results.json |
Per-seed held-out errors and the ablation comparison |
observations.npy, predictions.npy |
A replay input and verified inference output |
Evaluate or use the exported model without rerunning training:
cnskit evaluate --policy runs/malecns-task/policy --data runs/malecns-task/test.npz --objective regression
cnskit predict --policy runs/malecns-task/policy --input runs/malecns-task/observations.npy --out runs/replayed.npyThe example checks exact checkpoint replay and chunked stateful inference. Evaluation processes at most 1,024 timesteps per model call by default; use --chunk-size 256 for smaller prediction buffers. Evaluation dataset arrays and the graph still reside in memory. cnskit predict instead memory-maps NPY inputs and writes predictions incrementally; --chunk-size controls its time window. It processes one episode at a time and publishes output only on success. See file-backed inference for filesystem requirements. Output paths must be new. It loads the full prepared graph before selecting a subgraph; selection does not eliminate the import's memory requirement.
Supply observations [episode, time, input], regression targets [episode, time, output], unique string episode_ids, and an optional boolean mask in each NPZ split. Masked labels may be missing; all observations must remain finite. Split by independent sessions or environments rather than adjacent frames.
python examples/malecns_task.py --graph /data/traced-graph --out runs/my-task --train train.npz --validation validation.npz --test test.npzInput/output dimensions are inferred from your data. The default neuron selection is an engineering starting point, not a validated sensory-to-motor pathway. Use the exported task.json and Python API to define a task-relevant circuit, named channels and adaptation mode. For classification, see the classification example and training guide.
- Numerical correctness: five full-graph forward steps on 165,122 neurons and 25,563,197 edges matched an independent SciPy calculation to 5.96e-8 maximum absolute error.
- Software verification: 47 tests cover training, masking, early stopping, explicit-state inference, CLI evaluation and checkpoint integrity; CI runs Python 3.11 and 3.13.
- Task value: the included temporal-filter experiments test the workflow. They do not establish a MaleCNS advantage; the published synthetic-graph comparison favors the zero-edge ablation.
See measurements and reproducible commands. Full-graph differentiation, CUDA throughput, optimizer-state resume and biological behavior validation remain open work. The SDK is suitable for controlled research experiments; it is not yet a production training platform.
cnskit is the training API and CLI. The distribution name cns-tinker and existing cns_tinker imports remain for compatibility. This is a source release, not a claim of a published PyPI package. Legacy scenario APIs and synthetic game recipes remain available; their configuration-only cloud job interfaces are separate from the new local trainer.
Apache-2.0. Data is downloaded separately under upstream terms. See licensing.
Bring a task, a reproducible baseline, or an execution improvement. See contributing, validation and the roadmap. The most useful next result is a well-controlled task study that establishes when anatomical structure helps.