Skip to content
View yhyang201's full-sized avatar

Highlights

  • Pro

Block or report yhyang201

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

A curated resource list for learning AI performance engineering, from GPU fundamentals to production inference.

Python 3,762 346 Updated Sep 12, 2026

MAGI-2-preview: Scaling Video Generation Models Efficiently

Python 624 20 Updated Aug 6, 2026

Breakable CUDA graph capture for PyTorch - capture compatible regions as CUDA graphs while running incompatible operations eagerly between them.

Python 24 Updated Sep 15, 2026

My blog website.

JavaScript 525 40 Updated Sep 22, 2026

MoonEP: A Perfectly Balanced Expert Parallelism Library via Dynamic Redundant Experts

Python 1,151 134 Updated Sep 20, 2026

A tutorial on modern GPU programming for machine learning systems

HTML 1,293 139 Updated Sep 3, 2026

SGLang-native serving for the Moet sign-symmetric W2 expert format with SM120 W2/W4 kernels, GLM-5.2 NVFP4 TP4 on 4x RTX PRO 6000

Sass 20 Updated Jul 10, 2026

High-performance GPU kernels written in TIRx.

Python 104 13 Updated Sep 23, 2026

Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2

Rust 711 108 Updated Sep 23, 2026

Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond

Python 1,484 173 Updated Sep 23, 2026

A Python-embedded DSL that makes it easy to write fast, scalable ML kernels with minimal boilerplate.

Python 951 174 Updated Sep 23, 2026

A compiler, optimizer and executor for financial expressions and factors

C++ 322 62 Updated May 29, 2026

Kernel Design Agents (KDA) is a agent-centric workflow to write high-performance CUDA Kernels.

1,081 98 Updated Sep 14, 2026
Python 463 57 Updated Aug 26, 2026
Python 290 39 Updated Sep 5, 2026

Graphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini C…

TypeScript 83,858 7,068 Updated Sep 12, 2026

LLM KV cache compression made easy

Python 1,213 180 Updated Sep 21, 2026

Simple samples for TensorRT programming

Python 1,672 350 Updated Sep 13, 2026

TokenSpeed is a speed-of-light LLM inference engine.

Python 2,165 291 Updated Sep 23, 2026

A project to improve skills of large language models

Python 1,042 204 Updated Sep 23, 2026

high-performance linear attention kernel library built on TileLang

Python 706 79 Updated Sep 18, 2026

From Automated Idea Factory to Realization

Shell 1,458 128 Updated Aug 28, 2026

A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downs…

Python 3,861 620 Updated Sep 23, 2026

Run your GitHub Actions locally 🚀

Go 72,091 2,041 Updated Aug 9, 2026

CUDA kernels for linear attention variants, written in CuTe DSL and CUTLASS C++.

Python 551 69 Updated Sep 20, 2026

An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.

Rust 195,285 108,423 Updated Aug 16, 2026

An agentic skills framework & software development methodology that works.

Shell 290,650 26,007 Updated Sep 22, 2026

Bash is all you need - A nano claude code–like 「agent harness」, built from 0 to 1

Python 77,515 12,468 Updated Aug 26, 2026

A plug-and-play compiler that delivers free-lunch optimizations for both inference and training.

Python 333 29 Updated Sep 23, 2026
Next