Stars
Determined is an open-source machine learning platform that simplifies distributed training, hyperparameter tuning, experiment tracking, and resource management. Works with PyTorch and TensorFlow.
The core library and APIs implementing the Triton Inference Server.
A 10-week, 30-minutes-a-day roadmap for LLM inference serving and optimization. vLLM, SGLang, quantization, speculative decoding, benchmarking.
MSLK (Meta Superintelligence Labs Kernels) is a collection of PyTorch GPU operator libraries that are designed and optimized for GenAI training and inference, such as FP8 row-wise quantization and …
The Triton Inference Server provides an optimized cloud and edge inferencing solution.
A low-latency & high-throughput serving engine for LLMs
vLLM’s reference system for K8S-native cluster-wide deployment with community-driven performance optimization
Everything you need to know about LLM inference
TabFM (Tabular Foundation Model) is a pretrained tabular foundation model developed by Google Research for tabular data regression and classification.
Making large AI models cheaper, faster and more accessible
Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, …
20+ high-performance LLMs with recipes to pretrain, finetune and deploy at scale.
Tutel MoE: Optimized Mixture-of-Experts Library, Support GptOss/DeepSeek/Kimi-K2/Qwen3 using FP8/NVFP4/MXFP4
[TMLR 2026] LibMoE: A LIBRARY FOR COMPREHENSIVE BENCHMARKING MIXTURE OF EXPERTS IN LARGE LANGUAGE MODELS
Official repository for paper Routing-Free Mixture-of-Experts.
A framework for few-shot evaluation of language models.
wolfecameron / nanoMoE
Forked from karpathy/nanoGPTAn extension of the nanoGPT repository for training small MOE models.
Ongoing research training transformer models at scale
Pact improves Pi by retaining recent messages when compacting.
A PyTorch native platform for training generative AI models
PyTorch building blocks for the OLMo ecosystem
Autonomous experiment loop extension for pi
This library empowers users to seamlessly port pretrained models and checkpoints on the HuggingFace (HF) hub (developed using HF transformers library) into inference-ready formats that run efficien…
🐳 Efficient Triton implementations for "Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention"