A from-scratch Rust inference stack for AMD GPUs that talks straight to the Linux kernel driver. No ROCm, HIP, or HSA runtime. Serves OLMo 2 on MI355X.
-
Updated
Aug 15, 2026 - Rust
A from-scratch Rust inference stack for AMD GPUs that talks straight to the Linux kernel driver. No ROCm, HIP, or HSA runtime. Serves OLMo 2 on MI355X.
TMLR 2026 | Mechanistic interpretability: attention-head binding (EB*) as a marker of concept emergence. 7 models, 5 architectures (Pythia 160M–2.8B, OLMo-1B, CRFM GPT-2, SmolLM3-3B, Qwen2.5-1.5B), 41 terms.
Cross-Family Convergence of Neural Network Weight Skeletons. Companion to Zenodo paper (10.5281/zenodo.19652706).
OMTR — can LLM memorization be separated from predictability? Preregistered causal probing (activation patching) in Pythia & OLMo on consumer hardware; three honest non-separations, full corrections history
PyTorch DDP and multi-GPU LLM training — from a minimal distributed example to NanoGPT speedruns, Muon, H100 profiling, and modded-nanogpt. FBA LAB https://bubblnet.com
Pre-training a ~150M parameter code-specialized language model using OLMo 3 architecture (GQA, SWA, SwiGLU, RoPE) on PHP/JS/Python/C source code.
First open-source descriptor-augmented LLM for Neglected Tropical Disease drug discovery | OLMo-7B + QLoRA + DeepChem | Bioactivity & Toxicity prediction for Leishmaniasis, Chagas, Malaria, TB
OpenEuroLLM snapshot of the Berkeley Function Calling Leaderboard evaluation harness and OLMo evaluation orchestration.
RDKit-Guided Topological State Machine (TSM) for constrained SMILES generation with OLMo-7B. Solves BPE-tokenizer mismatch via RDKit-in-the-loop decoding. Achieved 100% validity on allenai/OLMo-7B-hf.
Paired OLMo continual-training experiment on spaced review and delayed FictionalQA retention
OpenEuroLLM tau2-bench evaluation harness with OLMo and Qwen serving, user-simulator, and aggregation scripts.
AI工具类(Midjourney / Notion / ChatGPT) 订阅教程类(Netflix / Spotify / Adobe) 虚拟卡支付类(Namecheap / OpenAI / Google)
Research workspace for model diffing between pretrained and post-trained language models.
Run LLM inference directly on AMD GPUs by bypassing ROCm for low-latency serving.
To associate your repository with the olmo topic, visit your repo's landing page and select "manage topics."