-
MLIR-iree Public
Forked from iree-org/ireeA retargetable MLIR-based machine learning compiler and runtime toolkit.
C++ Apache License 2.0 UpdatedMar 9, 2026 -
speculative-speculative-decoding Public
Forked from tanishqkumar/ssdA lightweight inference engine supporting speculative speculative decoding (SSD).
Python MIT License UpdatedMar 5, 2026 -
iree-amd-aie Public
Forked from nod-ai/iree-amd-aieIREE plugin repository for the AMD AIE accelerator
MLIR Apache License 2.0 UpdatedMar 4, 2026 -
hexagon-mlir Public
Forked from qualcomm/hexagon-mlirHexagon-MLIR is a compiler toolchain for compiling and executing AI kernels and models on Qualcomm Hexagon Neural Processing Units (NPUs).
C++ BSD 3-Clause "New" or "Revised" License UpdatedFeb 20, 2026 -
openpi Public
Forked from Physical-Intelligence/openpiPython Apache License 2.0 UpdatedDec 27, 2025 -
TensorRT-Model-Optimizer Public
Forked from NVIDIA/Model-OptimizerA unified library of state-of-the-art model optimization techniques like quantization, pruning, distillation, speculative decoding, etc. It compresses deep learning models for downstream deployment…
Python Apache License 2.0 UpdatedSep 10, 2025 -
EAGLE Public
Forked from SafeAILab/EAGLEOfficial Implementation of EAGLE-1 (ICML'24), EAGLE-2 (EMNLP'24), and EAGLE-3.
Python Apache License 2.0 UpdatedAug 13, 2025 -
transformers Public
Forked from huggingface/transformers🤗 Transformers: State-of-the-art Machine Learning for Pytorch, TensorFlow, and JAX.
Python Apache License 2.0 UpdatedAug 13, 2025 -
llvm-project-1 Public
Forked from llvm/llvm-projectThe LLVM Project is a collection of modular and reusable compiler and toolchain technologies.
LLVM Other UpdatedJul 24, 2025 -
SageAttention Public
Forked from thu-ml/SageAttentionQuantized Attention achieves speedup of 2-5x and 3-11x compared to FlashAttention and xformers, without lossing end-to-end metrics across language, image, and video models.
Cuda Apache License 2.0 UpdatedJul 21, 2025 -
pytorch Public
Forked from pytorch/pytorchTensors and Dynamic neural networks in Python with strong GPU acceleration
Python Other UpdatedJun 23, 2025 -
nncf Public
Forked from openvinotoolkit/nncfNeural Network Compression Framework for enhanced OpenVINO™ inference
Python Apache License 2.0 UpdatedApr 22, 2025 -
neural-compressor Public
Forked from intel/neural-compressorSOTA low-bit LLM quantization (INT8/FP8/INT4/FP4/NF4) & sparsity; leading model compression techniques on TensorFlow, PyTorch, and ONNX Runtime
Python Apache License 2.0 UpdatedApr 18, 2025 -
bitsandbytes Public
Forked from bitsandbytes-foundation/bitsandbytesAccessible large language models via k-bit quantization for PyTorch.
Python MIT License UpdatedMar 27, 2025 -
vllm Public
Forked from vllm-project/vllmA high-throughput and memory-efficient inference and serving engine for LLMs
Python Apache License 2.0 UpdatedMar 26, 2025 -
vision_transformer Public
Forked from google-research/vision_transformerJupyter Notebook Apache License 2.0 UpdatedMar 20, 2025 -
-
AutoAWQ Public
Forked from casper-hansen/AutoAWQAutoAWQ implements the AWQ algorithm for 4-bit quantization with a 2x speedup during inference. Documentation:
Python MIT License UpdatedMar 6, 2025 -
LookaheadDecoding Public
Forked from hao-ai-lab/LookaheadDecoding[ICML 2024] Break the Sequential Dependency of LLM Inference Using Lookahead Decoding
Python Apache License 2.0 UpdatedMar 6, 2025 -
-
Aria-MOE Public
Forked from rhymes-ai/AriaCodebase for Aria - an Open Multimodal Native MoE
Jupyter Notebook Apache License 2.0 UpdatedJan 22, 2025 -
titans-pytorch Public
Forked from lucidrains/titans-pytorchUnofficial implementation of Titans, SOTA memory for transformers, in Pytorch
Python MIT License UpdatedJan 15, 2025 -
Cosmos Public
Forked from NVIDIA/cosmosCosmos is a world model development platform that consists of world foundation models, tokenizers and video processing pipeline to accelerate the development of Physical AI at Robotics & AV labs. C…
Python Apache License 2.0 UpdatedJan 8, 2025 -
OpenStereo Public
Forked from XiandaGuo/OpenStereoOpenStereo: A Comprehensive Benchmark for Stereo Matching and Strong Baseline
Python UpdatedDec 27, 2024 -
Genesis Public
Forked from Genesis-Embodied-AI/genesis-worldA generative world for general-purpose robotics & embodied AI learning.
Python Apache License 2.0 UpdatedDec 20, 2024 -
blt Public
Forked from facebookresearch/bltCode for BLT research paper
Python BSD 3-Clause "New" or "Revised" License UpdatedDec 12, 2024 -
TensorRT-LLM Public
Forked from NVIDIA/TensorRT-LLMTensorRT-LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and build TensorRT engines that contain state-of-the-art optimizations to perform inference efficie…
C++ Apache License 2.0 UpdatedNov 12, 2024 -
deepcompressor Public
Forked from nunchux-ai/deepcompressorModel Compression Toolbox for Large Language Models and Diffusion Models
Python Apache License 2.0 UpdatedNov 10, 2024 -
-
OmniQuant Public
Forked from OpenGVLab/OmniQuant[ICLR2024 spotlight] OmniQuant is a simple and powerful quantization technique for LLMs.
Python MIT License UpdatedOct 8, 2024