Skip to content
View yhyang201's full-sized avatar
  • RadixArk

Highlights

  • Pro

Block or report yhyang201

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

My blog website.

JavaScript 461 34 Updated Aug 15, 2026

MoonEP: A Perfectly Balanced Expert Parallelism Library via Dynamic Redundant Experts

Python 1,079 119 Updated Aug 13, 2026

A tutorial on modern GPU programming for machine learning systems

HTML 1,163 125 Updated Aug 15, 2026

SGLang-native serving for the Moet sign-symmetric W2 expert format with SM120 W2/W4 kernels, GLM-5.2 NVFP4 TP4 on 4x RTX PRO 6000

Sass 20 Updated Jul 10, 2026

High-performance GPU kernels written in TIRx.

Python 87 7 Updated Aug 16, 2026

Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2

Rust 650 97 Updated Aug 16, 2026

Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond

Python 1,134 131 Updated Aug 13, 2026

A Python-embedded DSL that makes it easy to write fast, scalable ML kernels with minimal boilerplate.

Python 921 166 Updated Aug 15, 2026

A compiler, optimizer and executor for financial expressions and factors

C++ 316 60 Updated May 29, 2026
Python 377 46 Updated Aug 12, 2026
Python 247 35 Updated Aug 13, 2026

Graphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini C…

TypeScript 79,412 6,671 Updated Aug 11, 2026

LLM KV cache compression made easy

Python 1,173 168 Updated Aug 10, 2026

Simple samples for TensorRT programming

Python 1,664 350 Updated Jul 21, 2026

TokenSpeed is a speed-of-light LLM inference engine.

Python 1,905 242 Updated Aug 16, 2026

A project to improve skills of large language models

Python 1,024 194 Updated Aug 14, 2026

high-performance linear attention kernel library built on TileLang

Python 638 66 Updated Aug 11, 2026

From Automated Idea Factory to Realization

Shell 1,389 124 Updated Jul 18, 2026

A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downs…

Python 3,444 546 Updated Aug 16, 2026

Run your GitHub Actions locally 🚀

Go 71,507 2,001 Updated Aug 9, 2026

CUDA kernels for linear attention variants, written in CuTe DSL and CUTLASS C++.

Python 538 69 Updated Aug 14, 2026

An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.

Rust 195,053 109,111 Updated Aug 6, 2026

An agentic skills framework & software development methodology that works.

Shell 272,507 24,365 Updated Aug 13, 2026

Bash is all you need - A nano claude code–like 「agent harness」, built from 0 to 1

Python 74,311 12,033 Updated Aug 15, 2026

A plug-and-play compiler that delivers free-lunch optimizations for both inference and training.

Python 326 27 Updated Aug 14, 2026

Train speculative decoding models effortlessly and port them smoothly to SGLang serving.

Python 1,081 315 Updated Aug 14, 2026

An LLM-free Multi-dimensional Benchmark for Multi-modal Hallucination Evaluation

Python 172 7 Updated Jan 15, 2024

AI agents running research on single-GPU nanochat training automatically

Python 93,907 13,312 Updated Mar 26, 2026
Next