Skip to content
View ai-hpc's full-sized avatar
🦙
Research Physical AI agent safety, performance, memory
🦙
Research Physical AI agent safety, performance, memory

Highlights

  • Pro

Organizations

@openclaw @FastCrest @gittensor-model-hub

Block or report ai-hpc

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers.

Python 1,537 156 Updated Jul 24, 2026

A framework for few-shot evaluation of language models.

Python 13,408 3,440 Updated Jul 13, 2026
Python 50 224 Updated Jul 24, 2026

Fastest MoE/LLM inference runtime for consumer and edge Blackwell GPUs. SN74 on Gittensor.

Cuda 10 60 Updated Jul 24, 2026

SDK for developing enclaves

C 1,199 381 Updated Jul 24, 2026

Go ahead and axolotl questions

Python 12,248 1,398 Updated Jul 24, 2026
Python 38 30 Updated Jul 24, 2026

CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation

Python 1,119 93 Updated Jul 8, 2026

An optimized quantization and inference library for running LLMs locally on modern consumer-class GPUs

Python 1,065 123 Updated Jul 25, 2026

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python 10,881 1,616 Updated Jul 25, 2026

SGLang is a high-performance serving framework for large language models and multimodal models.

Python 30,736 7,396 Updated Jul 25, 2026

NanoCluster: Compact & Affordable Cluster for Everyone

222 9 Updated Jan 26, 2026

compiler learning resources collect.

Python 2,758 369 Updated May 20, 2026
Python 314 38 Updated Jun 9, 2026

[MLSys 2024 Best Paper Award] AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration

Python 3,595 316 Updated Jul 17, 2025

Real-Time VLAs via Future-state-aware Asynchronous Inference.

Python 439 32 Updated Apr 22, 2026

DFlash: Block Diffusion for Flash Speculative Decoding

Python 5,526 396 Updated May 10, 2026

CODA: Rewriting Transformer Blocks as GEMM-Epilogue Programs

Python 235 24 Updated Jul 25, 2026

An Optimizer for Nvidia Compilers.

Python 111 11 Updated Jul 3, 2026

Daily ArXiv Papers.

Python 446 103 Updated Jul 23, 2026

A retargetable MLIR-based machine learning compiler and runtime toolkit.

C++ 3,853 960 Updated Jul 25, 2026

Development repository for the Triton language and compiler

MLIR 19,782 3,044 Updated Jul 25, 2026

TokenSpeed is a speed-of-light LLM inference engine.

Python 1,658 199 Updated Jul 25, 2026

FlyDSL is the Python front‑end of the project: Flexible LaYout DSL.

Python 249 97 Updated Jul 25, 2026

High-performance, light-weight C++ LLM and VLM Inference Software for Physical AI

Python 483 93 Updated Jul 23, 2026

Medusa: Simple Framework for Accelerating LLM Generation with Multiple Decoding Heads

Jupyter Notebook 2,758 203 Updated Jun 25, 2024

Open Machine Learning Compiler Framework

Python 13,608 3,929 Updated Jul 25, 2026

All information and news with respect to Falcon-H1 series

122 15 Updated Oct 9, 2025

Model compression toolkit engineered for enhanced usability, comprehensiveness, and efficiency.

Python 1,485 168 Updated Jul 23, 2026

Interactive 3D visualization of dense decoder-only LLM inference. Companion to the AI Inference Engineer 2026 course.

TypeScript 1 Updated Jun 15, 2026
Next