Skip to content
View wesg52's full-sized avatar

Block or report wesg52

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

This repository collects all relevant resources about interpretability in LLMs

404 27 Updated Nov 1, 2024
Jupyter Notebook 9 1 Updated Mar 25, 2026

A curated list of LLM Interpretability related material - Tutorial, Library, Survey, Paper, Blog, etc..

308 13 Updated Jan 22, 2026

Training Sparse Autoencoders on Language Models

Python 1,499 258 Updated Aug 10, 2026

Create feature-centric and prompt-centric visualizations for sparse autoencoders (like those from Anthropic's published research).

HTML 268 48 Updated Feb 27, 2026

Stanford NLP Python library for understanding and improving PyTorch models via interventions

Python 897 109 Updated Mar 6, 2026

Investigating the generalization behavior of LM probes trained to predict truth labels: (1) from one annotator to another, and (2) from easy questions to hard

Python 33 5 Updated May 23, 2024

Code for my NeurIPS 2024 ATTRIB paper titled "Attribution Patching Outperforms Automated Circuit Discovery"

Jupyter Notebook 48 16 Updated May 31, 2024

Universal Neurons in GPT2 Language Models

Jupyter Notebook 30 7 Updated May 28, 2024

Representation Engineering: A Top-Down Approach to AI Transparency

Jupyter Notebook 1,020 132 Updated Aug 14, 2024

ModelDiff: A Framework for Comparing Learning Algorithms

Jupyter Notebook 60 4 Updated Aug 15, 2023

Contains random samples referenced in the paper "Sleeper Agents: Training Robustly Deceptive LLMs that Persist Through Safety Training".

150 28 Updated Mar 9, 2024

Mamba SSM architecture

Python 18,723 1,790 Updated Jul 22, 2026

The nnsight package enables interpreting and manipulating the internals of deep learned models.

Python 1,019 100 Updated Aug 10, 2026

Extracting spatial and temporal world models from LLMs

Jupyter Notebook 263 26 Updated Oct 17, 2023

Sparse Autoencoder for Mechanistic Interpretability

Python 303 44 Updated Jul 20, 2024

Tools for studying developmental interpretability in neural networks.

Python 150 32 Updated Apr 23, 2026

A Python package for interactive mapping and geospatial analysis with minimal coding in a Jupyter environment

Python 3,755 469 Updated Aug 8, 2026

A curated collection of resources and research related to the geometry of representations in the brain, deep networks, and beyond

1,076 74 Updated Feb 24, 2026

Compositional Linear Algebra

Python 518 34 Updated Aug 1, 2025
Jupyter Notebook 293 48 Updated Oct 1, 2024

Capture every activation and gradient of any PyTorch model — forward and backward — with automatic graph visualization, rich metadata, and live interventions. Works on any architecture, including d…

Python 654 29 Updated Jul 28, 2026

Interpretability for sequence generation models 🐛 🔍

Python 473 41 Updated Apr 25, 2026

Yet another redundant workflow engine

Python 597 52 Updated Jul 17, 2026

Kedro is a toolbox for production-ready data science. It uses software engineering best practices to help you create data engineering and data science pipelines that are reproducible, maintainable,…

Python 10,949 1,068 Updated Aug 10, 2026

Sparse probing paper full code.

Jupyter Notebook 68 11 Updated Dec 17, 2023
Next