Highlights
Lists (3)
Sort Name ascending (A-Z)
Starred repositories
Combining Double Soft-Min Critics with Adaptive KL Thresholding for Alignmet
H100-Optimized Goal-Conditioned Contrastive Self-RL for ARC Puzzles
Physics-based language model: O(n log n) wave field attention with linear-wave content routing. V4.1 achieves PPL 543 on WikiText-2.
The best-benchmarked open-source AI memory system. And it's free.
romannekrasovaillm / qqr
Forked from Alibaba-NLP/qqrqqr is an RL training framework for open-ended agents.
James' cookbook of evaluations and finetuning experiments
AuON( Alternative Unit-norm momentum-updates by Normal- ized nonlinear scaling), a linear-time optimizer that achieves remarkable perfor- mance at linear time without producing semi-orthogonal matrβ¦
Quick illustration of how one can easily read books together with LLMs. It's great and I highly recommend it.
This Repo Contains Script To Fine Tune Open Source Models Using Unsloth by using UI with simple click and progress
VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
Classify the morphologies of distant galaxies with cnn-vit
[CVPR2025] Breaking the Low-Rank Dilemma of Linear Attention
Leo optimizer, variation of Muon that runs faster
GLU Attention provide nearly cost-free performance boost for transformers with a simple mechanism that applies Gated Linear Unit to the values in Attention.
Awesome resources on normalizing flows.
ryyzn9 / R-Zero
Forked from Chengsong-Huang/R-Zerocodes for R-Zero: Self-Evolving Reasoning LLM from Zero Data (https://www.arxiv.org/pdf/2508.05004)
Stella Nera is the first Maddness accelerator achieving 15x higher area efficiency (GMAC/s/mm^2) and 25x higher energy efficiency (TMAC/s/W) than direct MatMul accelerators in the same technology
Continuous Thought Machines, because thought takes time and reasoning is a process.
American Sign Language Recognition System that translates Signs into their respective alphabets in real time
π€ LeRobot: Making AI for Robotics more accessible with end-to-end learning
Lecture notes on the RL series provided by Stanford.
This is the homepage of a new book entitled "Mathematical Foundations of Reinforcement Learning."