Stars
A framework for efficient model inference with omni-modality models
Mirage Persistent Kernel: Compiling LLMs into a MegaKernel
Implementation for FP8/INT8 Rollout for RL training without performence drop.
Production-tested AI infrastructure tools for efficient AGI development and community-driven innovation
ROCm / vllm
Forked from vllm-project/vllmA high-throughput and memory-efficient inference and serving engine for LLMs
2-2000x faster ML algos, 50% less memory usage, works on all hardware - new and old.
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
Disaggregated serving system for Large Language Models (LLMs).
Development repository for the Triton language and compiler
Accessible large language models via k-bit quantization for PyTorch.
Pax is a Jax-based machine learning framework for training large scale models. Pax allows for advanced and fully configurable experimentation and parallelization, and has demonstrated industry lead…
Distributed training framework for TensorFlow, Keras, PyTorch, and Apache MXNet.
A library for training and deploying machine learning models on Amazon SageMaker
Training and serving large-scale neural networks with auto parallelization.
Example pybind11 module built with a CMake-based build system
Seamless operability between C++11 and Python
Fast and memory-efficient exact attention
The Triton Inference Server provides an optimized cloud and edge inferencing solution.
Disruptive Behavior Mitigation Framework for Games