Anthropic Fellow working on interpretability, trying to figure out what is going on inside language models.
My most recent paper, Steering Awareness: Detecting Activation Steering from Within (COLM 2026), shows that models can be fine-tuned to detect and name steering vectors injected into their own activations.
Before research I spent several years as a software engineer. I studied Computer Science at the University of Chicago, specializing in Human-Computer Interaction, where I co-authored Touch&Fold (CHI 2021, Best Paper Honorable Mention).