Stars
Minimalistic 4D-parallelism distributed training framework for education purpose
fanshiqing / grouped_gemm
Forked from tgale96/grouped_gemmPyTorch bindings for CUTLASS grouped GEMM.
[ICLR 2026] When it comes to optimizers, it's always better to be safe than sorry
Get up and running with Kimi-K2.6, GLM-5.2, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.
Survey of Small Language Models from Penn State, ...
A subset of PyTorch's neural network modules, written in Python using OpenAI's Triton.
[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.
[ICLR 2025] DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
On-device AI across mobile, embedded and edge for PyTorch
LongLLaVA: Scaling Multi-modal LLMs to 1000 Images Efficiently via Hybrid Architecture
Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
Python Intelligence Config Manager. A superset of hydra+pydantic+lsp
FlagGems is an operator library for large language models implemented in the Triton Language.
Odysseus: Playground of LLM Sequence Parallelism
A framework for serving and evaluating LLM routers - save LLM costs without compromising quality
An Open Source Toolkit For LLM Distillation
A family of compressed models obtained via pruning and knowledge distillation
BitBLAS is a library to support mixed-precision matrix multiplications, especially for quantized LLM deployment.
Repository for Meta Chameleon, a mixed-modal early-fusion foundation model from FAIR.
The official evaluation suite and dynamic data release for MixEval.
source code of paper "On the Hallucination in Simultaneous Machine Translation"
[Neurips2024] Source code for xRAG: Extreme Context Compression for Retrieval-augmented Generation with One Token