Stars
A compact JEPA world model for goal-conditioned planning from pixels and actions — ICML 2026 workshop.
Official implementation of the paper "Next Embedding Prediction Makes World Models Stronger"
Official implementation of the paper "Enhancing Vision-Language Model Training with Reinforcement Learning in Synthetic Worlds for Real-World Success"
Code for the reproduction of counting manifolds
Official implementation of "Steering LLM Reasoning Through Bias-Only Adaptation" and "Small Vectors, Big Effects: A Mechanistic Study of RL-Induced Reasoning via Steering Vectors"
Official Triton kernels for TopK and HierarchicalTopK Sparse Autoencoder decoders.
Official implementation of the paper "You Do Not Fully Utilize Transformer's Representation Capacity"
Puzzles for learning Triton, play it with minimal environment configuration!
Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends
A simple, performant, and scalable Jax LLM!
[IEEE ICIP 2024] Diversifying Deep Ensembles: A Saliency Map Approach for Enhanced OOD Detection, Calibration, and Accuracy
🚀 Efficient implementations for emerging model architectures
Bend 2: a fast language that blocks AI mistakes via proof. Install: curl -fsSL https://bend-lang.com/install.sh | sh
Tile primitives for speedy kernels
GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection
Code for exploring Based models from "Simple linear attention language models balance the recall-throughput tradeoff"
Official implementation of the paper "Linear Transformers with Learnable Kernel Functions are Better In-Context Models"
A library for mechanistic interpretability of GPT-style language models
Code implementing "Efficient Parallelization of a Ubiquitious Sequential Computation" (Heinsen, 2023)
JAX-accelerated Meta-Reinforcement Learning Environments Inspired by XLand and MiniGrid 🏎️