I'm an engineer with research interests in AI safety, focused on detecting and measuring AI failures and catching them reliably.
Founding Engineer at Composo AI, building deployed LLM evaluation for high-stakes domains. I work on making automated evaluation reliable and practical: LLM-as-a-judge calibration, uncertainty quantification, classifier fine-tuning.
Previously SWE at Thought Machine, MSc (Machine Learning) at Imperial College London, working on Offline Reinforcement Learning.
- 📄 On Cost-Effective LLM-as-a-Judge Improvement Techniques (ICML 2026 Workshop “Statistical Frameworks for Uncertainty in Agentic Systems”; also accepted at ICML 2026 Workshop “Combining Theory and Benchmarks: Towards A Virtuous Cycle to Understand and Guarantee Foundation Model Performance”) - paper · code · blog
- 📝 Quantifying Blind Spots of LLM Evaluators - code · blog
- ✍️ More work at ryanlail.github.io