Skip to content
View yhyang201's full-sized avatar
  • RadixArk

Highlights

  • Pro

Block or report yhyang201

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

My blog website.

JavaScript 440 32 Updated Aug 8, 2026

MoonEP: A Perfectly Balanced Expert Parallelism Library via Dynamic Redundant Experts

Python 1,054 114 Updated Aug 7, 2026

A tutorial on modern GPU programming for machine learning systems

HTML 1,145 124 Updated Aug 6, 2026

SGLang-native serving for the Moet sign-symmetric W2 expert format with SM120 W2/W4 kernels, GLM-5.2 NVFP4 TP4 on 4x RTX PRO 6000

Sass 20 Updated Jul 10, 2026

High-performance GPU kernels written in TIRx.

Python 82 7 Updated Aug 9, 2026

Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2

Rust 634 97 Updated Aug 8, 2026

Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond

Python 1,130 129 Updated Aug 8, 2026

A Python-embedded DSL that makes it easy to write fast, scalable ML kernels with minimal boilerplate.

Python 919 161 Updated Aug 10, 2026

A compiler, optimizer and executor for financial expressions and factors

C++ 315 60 Updated May 29, 2026
Python 364 46 Updated Jun 9, 2026
Python 242 34 Updated Jul 29, 2026

Graphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini C…

TypeScript 78,602 6,599 Updated Jul 30, 2026

LLM KV cache compression made easy

Python 1,163 168 Updated Aug 6, 2026

Simple samples for TensorRT programming

Python 1,663 350 Updated Jul 21, 2026

TokenSpeed is a speed-of-light LLM inference engine.

Python 1,838 226 Updated Aug 10, 2026

A project to improve skills of large language models

Python 1,019 194 Updated Aug 8, 2026

high-performance linear attention kernel library built on TileLang

Python 628 65 Updated Aug 7, 2026

From Automated Idea Factory to Realization

Shell 1,379 123 Updated Jul 18, 2026

A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downs…

Python 3,412 536 Updated Aug 10, 2026

Run your GitHub Actions locally 🚀

Go 71,430 2,000 Updated Aug 9, 2026

CUDA kernels for linear attention variants, written in CuTe DSL and CUTLASS C++.

Python 535 70 Updated Aug 9, 2026

An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.

Rust 195,016 109,231 Updated Aug 6, 2026

An agentic skills framework & software development methodology that works.

Shell 269,783 24,121 Updated Aug 8, 2026

Bash is all you need - A nano claude code–like 「agent harness」, built from 0 to 1

Python 73,679 11,946 Updated Jul 28, 2026

A plug-and-play compiler that delivers free-lunch optimizations for both inference and training.

Python 325 28 Updated Aug 9, 2026

Train speculative decoding models effortlessly and port them smoothly to SGLang serving.

Python 1,058 311 Updated Aug 10, 2026

An LLM-free Multi-dimensional Benchmark for Multi-modal Hallucination Evaluation

Python 172 7 Updated Jan 15, 2024

AI agents running research on single-GPU nanochat training automatically

Python 93,526 13,286 Updated Mar 26, 2026
Next