Stars
Lightweight coding agent that runs in your terminal
DeepSeek Harness: Everything is a Plugin.
This is the homepage of a new book entitled "Mathematical Foundations of Reinforcement Learning."
A framework for efficient model inference with omni-modality models
Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels
Wan: Open and Advanced Large-Scale Video Generative Models
Valley is a cutting-edge multimodal large model designed to handle a variety of tasks involving text, images, video, and audio data.
A flexible and efficient training framework for large-scale alignment tasks
PyTorch distributed training acceleration framework
Qwen3 is the large language model series developed by Qwen team, Alibaba Cloud.
Efficient and easy multi-instance LLM serving
Artifact of OSDI '24 paper, ”Llumnix: Dynamic Scheduling for Large Language Model Serving“
DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
BitBLAS is a library to support mixed-precision matrix multiplications, especially for quantized LLM deployment.
A high-throughput and memory-efficient inference and serving engine for LLMs
The official repo of Pai-Megatron-Patch for LLM & VLM large scale training developed by Alibaba Cloud.
📷 EasyPhoto | Your Smart AI Photo Generator.
Transformer related optimization, including BERT, GPT
The Triton Inference Server provides an optimized cloud and edge inferencing solution.
Ongoing research training transformer models at scale
Chinese version of CLIP which achieves Chinese cross-modal retrieval and representation generation.
The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while control…
AITemplate is a Python framework which renders neural network into high performance CUDA/HIP C++ code. Specialized for FP16 TensorCore (NVIDIA GPU) and MatrixCore (AMD GPU) inference.