Skip to content
View ai-hpc's full-sized avatar
🦙
Research Inference Optimization & Confidential Computing
🦙
Research Inference Optimization & Confidential Computing

Highlights

  • Pro

Organizations

@openclaw @gittensor-ai-lab @FastCrest @cathedralai @gittensor-model-hub

Block or report ai-hpc

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers.

Python 1,546 160 Updated Jul 31, 2026

A framework for few-shot evaluation of language models.

Python 13,497 3,456 Updated Jul 13, 2026
Python 50 225 Updated Jul 31, 2026

Fastest MoE/LLM inference runtime for consumer and edge Blackwell GPUs. SN74 on Gittensor.

Cuda 12 61 Updated Jul 30, 2026

SDK for developing enclaves

C 1,200 381 Updated Jul 24, 2026

Go ahead and axolotl questions

Python 12,295 1,397 Updated Aug 1, 2026
Python 38 30 Updated Jul 31, 2026

CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation

Python 1,124 94 Updated Jul 8, 2026

An optimized quantization and inference library for running LLMs locally on modern consumer-class GPUs

Python 1,103 127 Updated Jul 31, 2026

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python 10,982 1,646 Updated Aug 2, 2026

SGLang is a high-performance serving framework for large language models and multimodal models.

Python 31,083 7,586 Updated Aug 2, 2026

NanoCluster: Compact & Affordable Cluster for Everyone

222 9 Updated Jan 26, 2026

compiler learning resources collect.

Python 2,758 369 Updated May 20, 2026
Python 320 39 Updated Jun 9, 2026

[MLSys 2024 Best Paper Award] AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration

Python 3,599 315 Updated Jul 17, 2025

Real-Time VLAs via Future-state-aware Asynchronous Inference.

Python 464 33 Updated Apr 22, 2026

DFlash: Block Diffusion for Flash Speculative Decoding

Python 5,556 397 Updated May 10, 2026

CODA: Rewriting Transformer Blocks as GEMM-Epilogue Programs

Python 235 24 Updated Aug 2, 2026

An Optimizer for Nvidia Compilers.

Python 113 11 Updated Jul 31, 2026

Daily ArXiv Papers.

Python 445 101 Updated Jul 30, 2026

A retargetable MLIR-based machine learning compiler and runtime toolkit.

C++ 3,866 967 Updated Aug 2, 2026

Development repository for the Triton language and compiler

MLIR 19,835 3,069 Updated Aug 2, 2026

TokenSpeed is a speed-of-light LLM inference engine.

Python 1,785 213 Updated Aug 2, 2026

FlyDSL is the Python front‑end of the project: a Flexible Layout Python DSL for expressing tiling, partitioning, data movement, and kernel structure at a high level.

Python 258 102 Updated Aug 2, 2026

High-performance, light-weight C++ LLM and VLM Inference Software for Physical AI

Python 489 98 Updated Jul 31, 2026

Medusa: Simple Framework for Accelerating LLM Generation with Multiple Decoding Heads

Jupyter Notebook 2,762 204 Updated Jun 25, 2024

Open Machine Learning Compiler Framework

Python 13,637 3,937 Updated Aug 1, 2026

All information and news with respect to Falcon-H1 series

122 15 Updated Oct 9, 2025

Model compression toolkit engineered for enhanced usability, comprehensiveness, and efficiency.

Python 1,501 171 Updated Aug 1, 2026

Interactive 3D visualization of dense decoder-only LLM inference. Companion to the AI Inference Engineer 2026 course.

TypeScript 1 Updated Jun 15, 2026
Next