On a quest to understand intelligence and ensure that advanced AGI is safe and beneficial.

Satvik Golechha

Hi! I’m a research scientist at the UK AI Security Institute (AISI). I work on post-training, interpreting, and evaluating models for various forms of misalignment, such as reward hacking, eval awareness, and sandbagging. Here’s why:

I believe AGI can lead to extraordinary well-being if models are capable, efficient, and aligned (which is what I work on!). By understanding how AI models learn and how they think, we can make them safer, better-aligned, and more useful for the future of life. This informs my research focus.

Previously, I worked on RL for efficient multi-turn exploration at CHAI at UC Berkeley. I was also a scholar at MATS Research (twice), where I worked on deception, feature geometry, and mechanistic interpretability.

Before moving full-time to AI safety, I worked at Microsoft Research on language models. Prior to that, I was an associate research scientist at Wadhwani AI working on AI for Social Good and Healthcare.

Writing fiction and poetry along the way!

If you’d like to discuss research, collaborate, or just chat about something, drop me an email!

Research

My research aims to better understand and control intelligence (via its emergence and expression in neural networks). Thus, I work on alignment-relevant interpretability, evals, and reinforcement learning for frontier AI systems and agents. Here is some of my recent published work:

Satvik Golechha, Sid Black, Joseph Bloom

Satvik Golechha, Sid Black, Joseph Bloom

Vikram Natarajan, Devina Jain, Shivam Arora, Satvik Golechha, Joseph Bloom

Jordan Taylor, Sid Black, Dillon Bowen, Thomas Read, Satvik Golechha, Alex Z-M., Oliver M., Connor K., Kola A., Jacob M., Sam Marks, Chris Cundy, Joseph Bloom

Satvik Golechha, Adrià Garriga-Alonso

David Chanin, James W.S., Tomáš D., Hardik B., Satvik Golechha, Joseph Bloom

Aly Lidayan, Jakob Bjorner, Satvik Golechha, Kartik Goyal, Alane Suhr

Samuel Marks, Johannes Treutlein, . . ., Satvik Golechha, . . ., Evan Hubinger

Ishwar B. , Hasith V. , Greta K., Ronan A. , Satvik Golechha

Satvik Golechha, Lucius Bushnaq, Euan Ong, Neeraj Kayal, Nandi Schoots

Satvik Golechha, Maheep C., Joan V., Alessandro Abate, Nandi Schoots

Satvik Golechha

Satvik Golechha

Satvik Golechha, James Dao

Pragya Srivastava*, Satvik Golechha*, Amit Deshpande, Amit Sharma

Pragnya R.*, Bhuvan S.*, Satvik Golechha*, Mohit Jain, and others

Mihir Kulkarni*, Satvik Golechha*, Rishi R.*, Jithin S.*, Alpan Raval

Poetry

Writing metaphorical poetry allows a channel into emotions that could not have been expressed another way. Check out my poetry page!

Almost done with my first poetry book, Anuswaad!

Fiction

A beautiful thing happens when fiction is written. A good story reflects back to us aspects of ourselves that we’re not aware of.

Really, it is the story that’s writing us.

Algebra to Zombies

A 29-week curriculum that covers most of the foundational math needed to do AI research. This accompanies a study group I used to run at Microsoft Research in India.

Research Blog

Some notes around AI research. For my research, please see my research statement and Scholar profile.

PS: For a more general (and hopefully fun) introduction to the less-taught parts of AI check out Alice!

Other Stuff

Intelligence: I write about intelligence and a number of related ideas in my fiction and research. I plan to bundle it into a blog series someday.

School: I’m writing a book (or a series of posts) on my version of an ideal school — I believe good schooling is highly impactful, undervalued, and achievable.

Like Winds & Dystop.ai: Slowly working on finishing these novels but aah so little time!

Infinite Jest: Reading this epic book; will take more than a year at my current pace.

London: I’ve moved to London and I’m looking to make friends, HMU!

All life is bound together by mutual support and interdependence.

Acharya UmaswaTi