Jacobian Lens (J-Lens) and J-Space Toolkit for transformer mechanistic interpretability. Train linear lenses, decompose hidden states, and run causal interventions on decoder-only models.
-
Updated
Sep 16, 2026 - Python
Jacobian Lens (J-Lens) and J-Space Toolkit for transformer mechanistic interpretability. Train linear lenses, decompose hidden states, and run causal interventions on decoder-only models.
An independent, from-scratch reproduction of the mechanistic-interpretability findings in Anthropic's When Models Manipulate Manifolds: The Geometry of a Counting Task
Mechanistic interpretability of multilingual reasoning in transformers. 170+ causal intervention experiments across 4 model families.
Code for SCIT: cache-level causal diagnostics for latent chain-of-thought models (EMNLP 2026 Findings)
Evaluate mechanistic estimates through the interventions, control loops, and safety decisions they guide.
Mechanistic interpretability of code models through probing, perturbations, and causal interventions
Mechanistic study of contextual-integrity post-training in Qwen2.5-7B, testing whether improved privacy behavior comes from new mechanisms or better use of machinery already present in the base model.
Experiments and analysis for studying overtopping in LLMs: high-leverage activation channels, intervention composition, dose and decode-time event localization, and learning-time causal reorganization.
A controlled mechanistic interpretability study testing whether representation–self-report alignment survives a targeted modification to a language model.
Behavioral auditing of finetuned language models using blinded policy recovery, model diffing, and causal activation interventions.
To associate your repository with the causal-interventions topic, visit your repo's landing page and select "manage topics."