- Evaluating spec-driven agentic development (Spec Kit vs. OpenSpec): where coding agents help and where humans need to stay in the loop
- Benchmarking AI coding agent configurations on reliability rate, tokens per success and turns to resolution
- Fine-tuning small local models with LoRA (mlx-lm) for a personal assistant
- Side builds: an iOS memory app and an ESP32 e-paper home dashboard
- Running coding agents in locked-down, sandboxed environments
- Planning and reasoning strategies in agentic systems
- LoRA fine-tuning and evaluation for small on-device models
- Measuring whether AI coding agents actually help, beyond vibes
- Multi-agent LLM systems (built one for real-time network optimization during my MSc thesis)
- Designing reproducible ML experiments
- Turning research prototypes into working tools
- Running local LLMs (Ollama, MLX) for practical tasks
- Python data workflows and ETL pipelines
Built my first LLM agent on a 5G testbed. Still not sure which part was harder.



