Skip to content
View ai-hpc's full-sized avatar
🦙
Research Inference Optimization & Confidential Computing
🦙
Research Inference Optimization & Confidential Computing

Highlights

  • Pro

Organizations

@openclaw @FastCrest @cathedralai @gittensor-model-hub

Block or report ai-hpc

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers.

Python 1,560 160 Updated Aug 11, 2026

A framework for few-shot evaluation of language models.

Python 13,591 3,473 Updated Aug 10, 2026
Python 50 227 Updated Aug 7, 2026

Fastest MoE/LLM inference runtime for consumer and edge Blackwell GPUs. SN74 on Gittensor.

Cuda 13 67 Updated Aug 11, 2026

SDK for developing enclaves

C 1,201 381 Updated Aug 10, 2026

Go ahead and axolotl questions

Python 12,335 1,401 Updated Aug 7, 2026
Python 38 30 Updated Aug 10, 2026

CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation

Python 1,128 95 Updated Jul 8, 2026

An optimized quantization and inference library for running LLMs locally on modern consumer-class GPUs

Python 1,127 127 Updated Aug 11, 2026

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python 11,100 1,686 Updated Aug 11, 2026

SGLang is a high-performance serving framework for large language models and multimodal models.

Python 31,654 7,789 Updated Aug 11, 2026

NanoCluster: Compact & Affordable Cluster for Everyone

222 8 Updated Jan 26, 2026

compiler learning resources collect.

Python 2,766 369 Updated May 20, 2026
Python 367 46 Updated Jun 9, 2026

[MLSys 2024 Best Paper Award] AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration

Python 3,612 318 Updated Jul 17, 2025

Real-Time VLAs via Future-state-aware Asynchronous Inference.

Python 476 34 Updated Apr 22, 2026

DFlash: Block Diffusion for Flash Speculative Decoding

Python 5,593 400 Updated May 10, 2026

CODA: Rewriting Transformer Blocks as GEMM-Epilogue Programs

Python 238 24 Updated Aug 6, 2026

An Optimizer for Nvidia Compilers.

Python 126 12 Updated Aug 6, 2026

Daily ArXiv Papers.

Python 446 101 Updated Aug 10, 2026

A retargetable MLIR-based machine learning compiler and runtime toolkit.

C++ 3,884 978 Updated Aug 10, 2026

Development repository for the Triton language and compiler

MLIR 19,921 3,096 Updated Aug 11, 2026

TokenSpeed is a speed-of-light LLM inference engine.

Python 1,843 231 Updated Aug 11, 2026

FlyDSL is the Python front‑end of the project: a Flexible Layout Python DSL for expressing tiling, partitioning, data movement, and kernel structure at a high level.

Python 260 104 Updated Aug 11, 2026

High-performance, light-weight C++ LLM and VLM Inference Software for Physical AI

Python 498 102 Updated Jul 31, 2026

Medusa: Simple Framework for Accelerating LLM Generation with Multiple Decoding Heads

Jupyter Notebook 2,764 205 Updated Jun 25, 2024

Open Machine Learning Compiler Framework

Python 13,665 3,948 Updated Aug 11, 2026

All information and news with respect to Falcon-H1 series

122 15 Updated Oct 9, 2025

Model compression toolkit engineered for enhanced usability, comprehensiveness, and efficiency.

Python 1,517 173 Updated Aug 7, 2026

Interactive 3D visualization of dense decoder-only LLM inference. Companion to the AI Inference Engineer 2026 course.

TypeScript 1 Updated Jun 15, 2026
Next