Skip to content
View fbarez's full-sized avatar
🌍
🌍

Highlights

  • Pro

Organizations

@torrvision @EdinAISafetyHub

Block or report fbarez

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

LLM Council works together to answer your hardest questions

Python 23,987 4,257 Updated Nov 22, 2025

Course Materials for Interpretability of Large Language Models (0368.4264) at Tel Aviv University

337 27 Updated Feb 8, 2026

Official code of "Rethinking Safety in LLM Fine-tuning: An Optimization Perspective" COLM 2025

Python 2 Updated Nov 14, 2025

This repository collects all relevant resources about interpretability in LLMs

404 27 Updated Nov 1, 2024

An open-source AI agent that brings the power of Gemini directly into your terminal.

TypeScript 106,491 14,428 Updated Aug 13, 2026

A curated list of resources for activation engineering

139 10 Updated Oct 2, 2025

Code and Data for Tau-Bench

Python 1,376 213 Updated Mar 18, 2026

Code for the EMNLP 2024 paper "Detecting and Mitigating Contextual Hallucinations in Large Language Models Using Only Attention Maps"

Python 152 12 Updated Oct 13, 2025

Simple, unified interface to multiple Generative AI providers

Python 16,087 1,696 Updated Jul 25, 2026

Solve Visual Understanding with Reinforced VLMs

Python 6,016 385 Updated Jul 7, 2026

An AI-powered research assistant that performs iterative, deep research on any topic by combining search engines, web scraping, and large language models. The goal of this repo is to provide the si…

TypeScript 19,544 1,990 Updated Apr 11, 2026

A curated list of LLM Interpretability related material - Tutorial, Library, Survey, Paper, Blog, etc..

308 13 Updated Jan 22, 2026

Benchmark to evaluate different LLMs for pragmatic (individual/context-specific) harms

Jupyter Notebook 4 3 Updated May 15, 2025

Sparse autoencoders

Python 1 Updated Oct 12, 2024

aider is AI pair programming in your terminal

Python 48,153 4,834 Updated May 22, 2026
Python 156 12 Updated Jul 21, 2026

LLM101n: Let's build a Storyteller

37,504 2,075 Updated Aug 1, 2024

HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Jupyter Notebook 1,029 153 Updated Aug 16, 2024

We focus on the behavior of AI, and the Cyber Soul. We investigate the alignment dynamics with deliberately designed experiments.

CSS 3 Updated Jul 14, 2024

Inspect: A framework for large language model evaluations

Python 2,530 651 Updated Aug 12, 2026

A trivial programmatic Llama 3 jailbreak. Sorry Zuck!

Python 576 65 Updated Jan 26, 2025

A collection of different ways to implement accessing and modifying internal model activations for LLMs

Jupyter Notebook 24 2 Updated Oct 18, 2024

Benchmark LLMs by fighting in Street Fighter 3! The new way to evaluate the quality of an LLM

Jupyter Notebook 1,484 182 Updated Mar 21, 2025

Evaluating LLMs with fewer examples

Jupyter Notebook 182 22 Updated Jul 4, 2026

Experimental AI Agents Framework

C# 306 61 Updated Jun 10, 2025

Explore what LLMs are really leanring over SFT

Python 28 1 Updated Mar 30, 2024
Next