I evaluate and improve AI systems for a living: I review how models generate code, catch where they hallucinate or break, and build the rubrics and verifiers that keep output quality honest. I write Python daily and reach for Java, C, Go, or TypeScript when the task calls for it. I recently reviewed frontier AI at Scale AI (Outlier) and I am now open to new roles.
My focus: find the exact turn where a model breaks, and build the deterministic check that proves it.
Three public tools that show one idea - evaluation as a hard gate, not a vibe - from three angles.
| Project | What it does | Links |
|---|---|---|
| TrajLens | Agentic-trajectory auditor. Grades tool-use traces against a 9-criterion deterministic rubric (including safety checks for buried errors and unauthorized actions) and measures an LLM-judge against human labels with Cohen's kappa. | Live - Code |
| JudgeLab | Measures GPT-4-as-a-judge against 3,355 real human expert votes from MT-Bench: Cohen's kappa 0.46, a 15.8% order-swap inconsistency rate, and quantified verbosity bias. | Live - Code |
| EvalGate | CI/CD pipeline that blocks bad models (GitHub Actions): tests, retrain, a hard ROC-AUC quality gate, artifact promotion, and an auto-deployed model dashboard. Running weekly, green for months. | Live - Code |
- AI Trainer and Reviewer - Scale AI (Outlier), 2026. Promoted from Attempter to Reviewer. Golden/silver agentic RL trajectory pipelines, rubric and pytest-verifier design, and safety-taxonomy grading (OpenClaw Atlas, Blue Shell, Lobster Safety). Confidential work, shown as process only.
- Software and ML Engineer - Tensium, 2026 to present (remote). CPU-based long-horizon ML tasks with multistep reasoning and code generation; develop, debug, and refactor Python pipelines for maintainability and performance.
- AI Trainer, LLM Evaluation - Handshake AI, 2026. Adversarial multi-hop question authoring plus an automated QC checker (Project Seal).
- MCX Advisor - Zebu Share and Wealth Management, 2025 to 2026. 4 years of trading experience across crypto and Indian equities; supported client trading operations and profit-and-loss updates.
Public process walkthroughs of the evaluation work (nothing confidential): pspmani.github.io/workflows.html
- Promoted from Attempter to Reviewer within 3 months at Scale AI (Outlier), for evaluation accuracy and consistent task quality.
- Shipped three live, public AI-evaluation tools: TrajLens, JudgeLab, and EvalGate.
- EvalGate CI/CD pipeline has run green for months, with automated weekly retraining and quality-gated deployments.
- Caught a session-level data-leakage bug and corrected an inflated 99% accuracy to a truthful 92.5% using GroupKFold.
- Vice President, Iterators Club - Jansons Institute of Technology (student leadership during B.E.).
- HandPilot - webcam gesture control for Windows: MediaPipe and OpenCV hand tracking for pointer, pinch clicks, scrolling, and an on-screen keyboard, with Windows CI tests.
- Chatbox - streaming chat app: FastAPI and Server-Sent Events backend, React and TypeScript frontend, pluggable model backend that runs offline with no API key.
- Customer Churn Predictor - Gradient Boosting on 7,043 records, 0.843 test ROC-AUC, SHAP explanations, deployed on Streamlit.
- Diabetes Foot Pressure Classification - caught session leakage with GroupKFold, correcting 99% to a realistic 92.5%.
- Spam Email Classifier - NLTK and TF-IDF pipeline, 97.8% accuracy with a Linear SVM.
- Claude Code 101 - Anthropic (2026)
- AI Fluency for Builders - Anthropic (2026)
- Introduction to Agents - Mercor (2026)
- Prompt Engineering (OutlierEDU) - Outlier AI (2026)
- Problem Solving (Intermediate) - HackerRank
- Frontend Developer (React) - HackerRank
- AWS Academy Cloud Foundations - Amazon Web Services (2022)
- Microsoft 365 Productivity Advanced - Microsoft (2022)
- Ethical Hacking and Penetration Testing - Scholiverse
- Enrolled Agent (EA) credential - in progress
Evaluation: LLM evaluation, RLHF / RLAIF, red teaming, prompt engineering, rubric design, agentic trajectory review, hallucination detection, multimodal assessment.
Also: US Enrolled Agent (EA) credential in progress; Indian income-tax returns (ITR-1 to ITR-4).
Open to AI Evaluation, LLM Engineering, RLHF / RLAIF, and ML Engineer roles.
Portfolio |
LinkedIn |
Email