Skip to content
View reiase's full-sized avatar

Block or report reiase

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Scaling the Horizon, Not the Parameters

Python 520 50 Updated Jul 16, 2026

Ultra Kernel Samepage Merging patches (aka uksm, ksm)

19 2 Updated Jan 7, 2025

Ultra-light Harness scaffolding for AI agents, a mini version of claude code

Python 934 358 Updated Jun 10, 2026

Fast, small, and fully autonomous AI personal assistant infrastructure, any OS, any platform — deploy anywhere, swap anything 🦀

Rust 32,372 4,836 Updated Jul 24, 2026

A security-focused library OS supporting kernel- and user-mode execution

Rust 2,651 134 Updated Jul 24, 2026

Awesome-LLM-KV-Cache: A curated list of 📙Awesome LLM KV Cache Papers with Codes.

460 29 Updated Jun 17, 2026

A minimal, secure Python interpreter written in Rust for use by AI

Rust 7,935 394 Updated Jul 22, 2026

Our first fully AI generated deep learning system

Python 632 48 Updated Feb 2, 2026

LightRFT (Light Reinforcement Fine-Tuning) is an advanced reinforcement learning fine-tuning framework designed for Large Language Models (LLMs) and Vision-Language Models (VLMs).

Python 19 3 Updated Jan 12, 2026

Serverless LLM Serving for Everyone.

Python 693 75 Updated May 4, 2026

GEAR: An Efficient KV Cache Compression Recipefor Near-Lossless Generative Inference of LLM

Python 183 20 Updated Jul 12, 2024

Nano vLLM

Python 14,624 2,352 Updated Apr 26, 2026

The absolute trainer to light up AI agents.

Python 17,421 1,524 Updated Jul 16, 2026

MSCCL++: A GPU-driven communication stack for scalable AI applications

C++ 542 102 Updated Jul 24, 2026

Awesome list for LLM quantization

Python 434 26 Updated Apr 20, 2026

ASTRA-sim2.0: Modeling Hierarchical Networks and Disaggregated Systems for Large-model Training at Scale

C++ 648 219 Updated Apr 25, 2026

High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale

Rust 5,656 527 Updated Jul 22, 2026

Dynamic Performance Profiler for Distributed AI

Rust 11 5 Updated Jul 18, 2026

InfiniGen: Efficient Generative Inference of Large Language Models with Dynamic KV Cache Management (OSDI'24)

Python 192 40 Updated Jul 10, 2024

Unified KV Cache Compression Methods for Auto-Regressive Models

Python 1,355 178 Updated Jul 10, 2026

[ICML 2025] Official PyTorch implementation of "FlatQuant: Flatness Matters for LLM Quantization"

Python 223 34 Updated Nov 25, 2025

This repository contains everything you need to become proficient in System Design

4,773 672 Updated Oct 6, 2024

linux-tkg custom kernels

Shell 1,589 201 Updated Jul 21, 2026
Rust 2 2 Updated Aug 21, 2025

PyTorch Single Controller

Rust 1,060 164 Updated Jul 24, 2026

收集与绘制火焰图(rust版)

Rust 4 3 Updated Jan 20, 2026

所有小初高、大学PDF教材。

Roff 76,078 17,125 Updated Oct 18, 2025
Python 89 15 Updated Apr 18, 2025

LLM KV cache compression made easy

Python 1,145 159 Updated Jul 9, 2026
Next