-
Harvard University
- Boston, MA
Stars
maze datasets for investigating OOD behavior of ML systems
Mobile and Web client for Codex and Claude Code, with realtime voice, encryption and fully featured
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks [ICLR 2025]
📊 A simple command-line utility for querying and monitoring GPU status
Stanford NLP Python library for understanding and improving PyTorch models via interventions
Stanford NLP Python library for benchmarking the utility of LLM interpretability methods
Real-time Claude Code usage monitor with predictions and warnings
A powerful GUI app and Toolkit for Claude Code - Create custom agents, manage interactive Claude Code sessions, run secure background agents, and more.
Steering vectors for transformer language models in Pytorch / Huggingface
Bayesian scaling laws for in-context learning.
Claude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining complex code, and handling git workflo…
TextWorld is a sandbox learning environment for the training and evaluation of reinforcement learning (RL) agents on text-based games.
A probabilistic programming language for metacognitive modeling
Steering Llama 2 with Contrastive Activation Addition
The nnsight package enables interpreting and manipulating the internals of deep learned models.
A high-throughput and memory-efficient inference and serving engine for LLMs
👨💻 An awesome and curated list of best code-LLM for research.
A library for mechanistic interpretability of GPT-style language models
Training Sparse Autoencoders on Language Models
Sparsify transformers with SAEs and transcoders
The hub for EleutherAI's work on interpretability and learning dynamics
Productive, portable, and performant GPU programming in Python.