Stars
LLM Council works together to answer your hardest questions
Course Materials for Interpretability of Large Language Models (0368.4264) at Tel Aviv University
Official code of "Rethinking Safety in LLM Fine-tuning: An Optimization Perspective" COLM 2025
This repository collects all relevant resources about interpretability in LLMs
An open-source AI agent that brings the power of Gemini directly into your terminal.
A curated list of resources for activation engineering
Code for the EMNLP 2024 paper "Detecting and Mitigating Contextual Hallucinations in Large Language Models Using Only Attention Maps"
Simple, unified interface to multiple Generative AI providers
Solve Visual Understanding with Reinforced VLMs
An AI-powered research assistant that performs iterative, deep research on any topic by combining search engines, web scraping, and large language models. The goal of this repo is to provide the si…
A curated list of LLM Interpretability related material - Tutorial, Library, Survey, Paper, Blog, etc..
Benchmark to evaluate different LLMs for pragmatic (individual/context-specific) harms
aider is AI pair programming in your terminal
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
We focus on the behavior of AI, and the Cyber Soul. We investigate the alignment dynamics with deliberately designed experiments.
Inspect: A framework for large language model evaluations
A trivial programmatic Llama 3 jailbreak. Sorry Zuck!
A collection of different ways to implement accessing and modifying internal model activations for LLMs
Benchmark LLMs by fighting in Street Fighter 3! The new way to evaluate the quality of an LLM
Evaluating LLMs with fewer examples
Explore what LLMs are really leanring over SFT