On a quest to understand intelligence and ensure that advanced AGI is safe and beneficial.
Hi! I’m a research scientist at the UK AI Security Institute (AISI). I work on post-training, interpreting, and evaluating models for various forms of misalignment, such as reward hacking, eval awareness, and sandbagging. Here’s why:
I believe AGI can lead to extraordinary well-being if models are capable, efficient, and aligned (which is what I work on!). By understanding how AI models learn and how they think, we can make them safer, better-aligned, and more useful for the future of life. This informs my research focus.
Previously, I worked on RL for efficient multi-turn exploration at CHAI at UC Berkeley. I was also a scholar at MATS Research (twice), where I worked on deception, feature geometry, and mechanistic interpretability.
Before moving full-time to AI safety, I worked at Microsoft Research on language models. Prior to that, I was an associate research scientist at Wadhwani AI working on AI for Social Good and Healthcare.
If you’d like to discuss research, collaborate, or just chat about something, drop me an email!
Research
My research aims to better understand and control intelligence (via its emergence and expression in neural networks). Thus, I work on alignment-relevant interpretability, evals, and reinforcement learning for frontier AI systems and agents. Here is some of my recent published work:
KL Penalties in RL can Increase CoT Unfaithfulness
2026, UK AISI
One Probe Won’t Catch Them All: Targeted Deception Detection
ICML 2026 (co-mentored at LASR Labs)
Auditing Games for Sandbagging
2025, UK AISI (with FAR AI)
Among Us: A Sandbox for Measuring and Detecting Agentic Deception
NeurIPS 2025 (Spotlight) (MATS)
A is for Absorption: Studying Feature Splitting and Absorption in SAEs
NeurIPS 2025 (Oral) (MATS)
ABBEL: Natural-Language Belief States for Memory-Efficient Interaction
NeurIPS 2025 (Spotlight at LAW workshop) (CHAI, UC Berkeley)
Auditing Language Models for Hidden Objectives
2025, Anthropic (external collaboration)
Who’s the Evil Twin? Differential Auditing for Undesired Behavior
ICML 2026 (AI4G workshop) (mentored at SPAR 2025)
Intricacies of Feature Geometry in Large Language Models
ICLR 2025 (MATS Research)
Studying Cross-cluster Modularity in Neural Networks
NeurIPS 2024 (SoDL workshop) (MATS Research)
Some Lessons from the OpenAI-FrontierMath Debacle
Satvik Golechha
Progress Measures for Grokking on Real-world Tasks
ICML 2024 (HiDL workshop) (independent)
Challenges in Mechanistically Interpreting Harmful Representations
ICML 2024 (MI workshop) (independent)
NICE: To Optimize In-Context Examples or Not?
ACL 2024 (Microsoft Research)
BYoEB: An LLM-Powered Expert-in-the-Loop Chat System
UbiComp 2025 (Microsoft Research)
Predicting Treatment Adherence of Tuberculosis Patients at Scale
NeurIPS 2022 (Wadhwani AI)
Poetry
Writing metaphorical poetry allows a channel into emotions that could not have been expressed another way. Check out my poetry page!
Fiction
A beautiful thing happens when fiction is written. A good story reflects back to us aspects of ourselves that we’re not aware of.
Really, it is the story that’s writing us.
Research Blog
Some notes around AI research. For my research, please see my research statement and Scholar profile.
PS: For a more general (and hopefully fun) introduction to the less-taught parts of AI check out Alice!
Other Stuff
Intelligence: I write about intelligence and a number of related ideas in my fiction and research. I plan to bundle it into a blog series someday.
School: I’m writing a book (or a series of posts) on my version of an ideal school — I believe good schooling is highly impactful, undervalued, and achievable.
Like Winds & Dystop.ai: Slowly working on finishing these novels but aah so little time!
Infinite Jest: Reading this epic book; will take more than a year at my current pace.
London: I’ve moved to London and I’m looking to make friends, HMU!