Stars
Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels
Maximally Scale-Stable Parameterization (MSSP) for Mixture-of-Experts: hyperparameter transfer and predictable scaling across width, depth, expert count, and expert width.
Solve puzzles. Improve your pytorch.
Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours
MoE training for Me and You and maybe other people
Automatic tiling window manager for macOS à la xmonad.
wolfecameron / nanoMoE
Forked from karpathy/nanoGPTAn extension of the nanoGPT repository for training small MOE models.
Continual pretrainig experiments with large language models
Ongoing research training transformer models at scale
Matplotlib styles for scientific plotting
Code for the Paper: "STAMP Your Content: Proving Dataset Membership via Watermarked Rephrasings"
Code release to accompany the paper "Persistent Pre-training Poisoning of LLMs"
[NeurIPS D&B '25] The one-stop repository for LLM unlearning
Open-source framework for the research and development of foundation models.
PyTorch building blocks for the OLMo ecosystem
Monitor the training of your PyTorch models, including transformers
The code for creating the iGSM datasets in papers "Physics of Language Models Part 2.1, Grade-School Math and the Hidden Reasoning Process" (arxiv 2407.20311) and "Physics of Language Models Part 2…
What's In My Big Data (WIMBD) - a toolkit for analyzing large text datasets
EleutherAI / nanoGPT-mup
Forked from karpathy/nanoGPTThe simplest, fastest repository for training/finetuning medium-sized GPTs.
A library for unit scaling in PyTorch