Skip to content
View catswe's full-sized avatar

Block or report catswe

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

An extension of the nanoGPT repository for training small MOE models.

Python 282 32 Updated Mar 9, 2025

Ongoing research training transformer models at scale

Python 17,242 4,294 Updated Jul 28, 2026

Pact improves Pi by retaining recent messages when compacting.

TypeScript 2 Updated Jul 1, 2026

A PyTorch native platform for training generative AI models

Python 5,569 914 Updated Jul 28, 2026

PyTorch building blocks for the OLMo ecosystem

Python 1,433 296 Updated Jul 28, 2026

Autonomous experiment loop extension for pi

TypeScript 7,274 429 Updated Jul 15, 2026

This library empowers users to seamlessly port pretrained models and checkpoints on the HuggingFace (HF) hub (developed using HF transformers library) into inference-ready formats that run efficien…

Python 93 92 Updated Jul 28, 2026

🐳 Efficient Triton implementations for "Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention"

Python 1,014 53 Updated Feb 5, 2026
Python 193 19 Updated Oct 31, 2025

NanoGPT (124M) in 90 seconds

Python 5,597 849 Updated Jul 28, 2026

Speedrunning LoRA fine-tuning: frozen task, frozen hardware, public wall-clock leaderboard. modded-nanogpt for fine-tuning.

Python 144 10 Updated Jul 27, 2026

Tiered optimizer state allocation for memory-efficient MoE training. Cuts optimizer memory by 97.4%, outperforming AdamW/Muon/Lion while fitting a 6.78B MoE on a single 40GB GPU.

Python 11 2 Updated Jul 22, 2026

Code for Adam-mini: Use Fewer Learning Rates To Gain More https://arxiv.org/abs/2406.16793

Python 457 19 Updated May 13, 2025

Code for the signSGD paper

Jupyter Notebook 95 16 Updated Jan 12, 2021

🦁 Lion, new optimizer discovered by Google Brain using genetic algorithms that is purportedly better than Adam(w), in Pytorch

Python 2,195 54 Updated Jul 9, 2026

TokenSpeed is a speed-of-light LLM inference engine.

Python 1,712 203 Updated Jul 28, 2026

Root Mean Square Layer Normalization

Python 283 19 Updated Mar 28, 2023

Repo for "Monarch Mixer: A Simple Sub-Quadratic GEMM-Based Architecture"

Assembly 564 45 Updated Dec 28, 2024

Building blocks for foundation models.

635 27 Updated Jan 3, 2024

[ICLR2025] Codebase for "ReMoE: Fully Differentiable Mixture-of-Experts with ReLU Routing", built on Megatron-LM.

Python 118 11 Updated Dec 20, 2024

Official code for UnICORNN (ICML 2021)

Python 28 3 Updated Oct 1, 2021

Achieve state of the art inference performance with modern accelerators on Kubernetes

Shell 3,897 638 Updated Jul 28, 2026

DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms

Python 6,800 632 Updated Jul 9, 2026

A dynamic binary instrumentation tool for tracing and analyzing CUDA kernel instructions.

Python 76 9 Updated Jul 27, 2026

🎬 3.7× faster video generation E2E 🖼️ 1.6× faster image generation E2E ⚡ ColumnSparseAttn 9.3× vs FlashAttn‑3 💨 ColumnSparseGEMM 2.5× vs cuBLAS

Cuda 111 2 Updated Sep 8, 2025

Causal depthwise conv1d in CUDA, with a PyTorch interface

Cuda 923 202 Updated May 9, 2026

Ongoing research training transformer language models at scale, including: BERT & GPT-2

Python 1,448 226 Updated Mar 20, 2024

[AAAI 2026] Official implementation of "FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models". If you find this repository helpful, please consider starring 🌟 it to support the p…

Python 17 2 Updated May 1, 2026

Official repository of the xLSTM.

Python 2,189 185 Updated May 28, 2026
Next