Skip to content
View Tracin's full-sized avatar

Block or report Tracin

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞

TypeScript 386,748 81,259 Updated Aug 19, 2026

A safetensors extension to efficiently store sparse quantized tensors on disk

Python 313 113 Updated Aug 19, 2026

Expert Parallelism Load Balancer

Python 1,422 208 Updated Mar 24, 2025

🙌 OpenHands: AI-Driven Development

TypeScript 84,460 10,995 Updated Aug 19, 2026

Modern CUDA Learn Notes with PyTorch for Beginners, 200+ CUDA Kernels, Tensor Cores, HGEMM, FA-2 MMA.

Cuda 11,779 1,239 Updated Aug 17, 2026

⚡️Write HGEMM from scratch using Tensor Cores with WMMA, MMA and CuTe API, Achieve Peak⚡️ Performance.

Cuda 158 10 Updated May 10, 2025

Structured Outputs

Python 15,647 862 Updated Aug 16, 2026

VPTQ, A Flexible and Extreme low-bit quantization algorithm

Python 680 52 Updated Aug 4, 2026

Fast Matrix Multiplications for Lookup Table-Quantized LLMs

C++ 394 20 Updated Apr 13, 2025

🐳 Efficient Triton implementations for "Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention"

Python 1,019 54 Updated Feb 5, 2026

DeepGEMM: clean and efficient BLAS kernel library on GPU

Cuda 7,703 1,183 Updated Aug 11, 2026

DeepEP: an efficient expert-parallel communication library

Cuda 10,030 1,394 Updated Aug 5, 2026

FlashMLA: Efficient Multi-head Latent Attention Kernels

C++ 12,853 1,126 Updated Jul 28, 2026

Train transformer language models with reinforcement learning.

Python 19,107 2,916 Updated Aug 19, 2026

High-Performance FP32 GEMM on CUDA devices

Cuda 126 8 Updated Jan 21, 2025

Code for Neurips24 paper: QuaRot, an end-to-end 4-bit inference of large language models.

Python 531 75 Updated Nov 26, 2024

Fast Hadamard transform in CUDA, with a PyTorch interface

C 345 70 Updated Mar 10, 2026

[EMNLP 2024 & AAAI 2026] A powerful toolkit for compressing large models including LLMs, VLMs, and video generative models.

Python 740 77 Updated May 14, 2026

📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉

Python 5,462 430 Updated Aug 14, 2026

Building a quick conversation-based search demo with Lepton AI.

TypeScript 8,074 993 Updated Dec 2, 2025

FP16xINT4 LLM inference kernel that can achieve near-ideal ~4x speedups up to medium batchsizes of 16-32 tokens.

Python 1,131 91 Updated Sep 4, 2024

Official implementation of Half-Quadratic Quantization (HQQ)

Python 953 94 Updated Feb 26, 2026

Built upon Megatron-Deepspeed and HuggingFace Trainer, EasyLLM has reorganized the code logic with a focus on usability. While enhancing usability, it also ensures training efficiency.

Python 48 8 Updated Sep 18, 2024

Code for paper: "QuIP: 2-Bit Quantization of Large Language Models With Guarantees"

Python 402 38 Updated Feb 24, 2024

Code and documents of LongLoRA and LongAlpaca (ICLR 2024 Oral)

Python 2,687 281 Updated Aug 14, 2024

The Triton TensorRT-LLM Backend

941 143 Updated Aug 17, 2026

Awesome LLM compression research papers and tools.

1,862 130 Updated Jun 30, 2026

TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. Tensor…

Python 14,418 2,674 Updated Aug 19, 2026

A more memory-efficient rewrite of the HF transformers implementation of Llama for use with quantized weights.

Python 2,938 220 Updated Sep 30, 2023

Offline Quantization Tools for Deploy.

Python 143 19 Updated Dec 28, 2023
Next