Hi, I’m Ryan, a Founding Engineer at Composo AI in London. My current work focuses on making automated evaluation reliable enough to deploy, through LLM-as-a-judge calibration, uncertainty quantification, and classifier fine-tuning for guardrails.
I use this site to write up papers I’ve worked on and think through research in progress.
Posts
-
Quantifying blind spots of LLM Evaluators
When an LLM judge is wrong, does its own resample disagreement know?
-
Making LLM judges more reliable without fine-tuning
Improving RewardBench 2 performance by 13.5pp; accepted at the ICML 2026 Workshop "Statistical Frameworks for Uncertainty in Agentic Systems"