Skip to content
View GaoYusong's full-sized avatar

Block or report GaoYusong

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Starred repositories

Showing results

👻 Ghostty is a fast, feature-rich, and cross-platform terminal emulator that uses platform-native UI and GPU acceleration.

Zig 61,444 3,494 Updated Sep 22, 2026

Meta-Framework of Spatiotemporal Composability

TypeScript 8,761 542 Updated Sep 8, 2026

DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm

C 22,648 2,168 Updated Sep 20, 2026

The fastest, litest AI Gateway. Rust core with Python SDK. Call 100+ LLM APIs in OpenAI (or native) format with cost tracking, guardrails, load balancing, and logging [Bedrock, Azure, OpenAI, Anthr…

Python 59,433 11,683 Updated Sep 23, 2026

TokenSpeed is a speed-of-light LLM inference engine.

Python 2,161 291 Updated Sep 23, 2026

Use Codex from Claude Code to review code or delegate tasks.

JavaScript 33,495 2,328 Updated Jul 8, 2026

A Claude Code plugin that shows what's happening - context usage, active tools, running agents, and todo progress

JavaScript 28,118 1,299 Updated Sep 19, 2026

Garry's Opinionated OpenClaw/Hermes Agent Brain

TypeScript 30,247 4,527 Updated Sep 23, 2026

Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models

Python 4,692 359 Updated Jan 14, 2026

A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.

Python 5,138 865 Updated May 17, 2026

Tile-Based Runtime for Ultra-Low-Latency LLM Inference

Python 1,793 125 Updated Aug 13, 2026

Open-source book with Modern CUDA Learn Notes for Beginners, includes FP16/BF16, FP8, HGEMM, FlashAttention, CuTe, etc.

Cuda 11,992 1,259 Updated Sep 23, 2026

slime is an LLM post-training framework for RL Scaling.

Python 8,524 1,267 Updated Sep 23, 2026

Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels

Python 7,468 750 Updated Sep 23, 2026

The source of LMSYS website and blogs

JavaScript 98 113 Updated Sep 21, 2026

A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downs…

Python 3,846 614 Updated Sep 23, 2026

A fast communication-overlapping library for tensor/expert parallelism on GPUs.

C++ 1,363 115 Updated Aug 28, 2025

Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, DeepSeek-V4, MiniMax-H3, Gemma 4, FLUX and more.

Python 76,604 7,013 Updated Sep 23, 2026

TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. Tensor…

Python 14,701 2,772 Updated Sep 23, 2026

Mirage Persistent Kernel: Compiling LLMs into a MegaKernel

Cuda 2,511 257 Updated Sep 20, 2026

DuckLake is an integrated data lake and catalog format

C++ 2,995 258 Updated Sep 22, 2026

Nano vLLM

Python 15,587 2,634 Updated Apr 26, 2026

Kernels, of the mega variety :)

Python 832 68 Updated May 26, 2026

Analyze computation-communication overlap in V3/R1.

1,188 152 Updated Mar 21, 2025

Documented system prompts from Anthropic - Claude Fable 5.1, Opus 5.5, Claude Design, Claude Code. OpenAI - ChatGPT GPT-6-Astra, Codex. Google - Gemini 3.8 Flash, 3.1 Pro, Antigravity. xAI - Grok, …

JavaScript 68,102 11,056 Updated Sep 22, 2026

Production-grade client-side tracing, profiling, and analysis for complex software systems.

C++ 6,542 876 Updated Sep 23, 2026

A high-throughput and memory-efficient inference and serving engine for LLMs

Python 92,466 22,543 Updated Sep 23, 2026

My learning notes for ML SYS.

HTML 7,393 510 Updated Sep 20, 2026

NVIDIA Inference Xfer Library (NIXL)

C++ 1,266 452 Updated Sep 23, 2026
Next