- Beijing
-
00:06
(UTC +08:00) - https://scholar.google.com/citations?hl=zh-CN&user=MBR97ZIAAAAJ
Stars
A vLLM patch + hand‑written SM120 SASS kernels: 2‑bit MoE experts + an FP4 "delta" cache that recovers precision — matching the official (NV)FP4 checkpoint's quality on consumer Blackwell cards
Reference implementation and examples of the CuTe Layout representation and algebra.
Orca is the ADE for working with a fleet of parallel agents. Run any coding agent with your own subscription. Available on desktop, mobile and VPS.
Bridge Feishu/Lark to AI coding CLIs — Claude Code, Codex, Gemini, OpenCode… every DM, group or topic spawns its own live-streaming CLI session
TurboQuant: Near-optimal KV cache quantization for LLM inference (3-bit keys, 2-bit values) with Triton kernels + vLLM integration
Persistent Claude/Codex terminal and Agent Workspace dashboard backed by ttyd + tmux.
TokenSpeed is a speed-of-light LLM inference engine.
A kernel library written in tilelang
Efficient and unified implementations for TopK-based sparse attention
An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.
CUDA Tile IR is an MLIR-based intermediate representation and compiler infrastructure for CUDA kernel optimization, focusing on tile-based computation patterns and optimizations targeting NVIDIA te…
Bridge local AI coding agents (Claude Code, Cursor, Gemini CLI, Codex) to messaging platforms (Feishu/Lark, DingTalk, Slack, Telegram, Discord, LINE, WeChat Work). Chat with your AI dev assistant f…
PTX ISA 9.1 documentation converted to searchable markdown. Includes Claude Code skill for CUDA development.
Agentic Kernel Optimization for All — automated GPU kernel optimization for any kernel, any hardware, any language
Automated CUDA kernel performance diagnostics from NVIDIA Nsight Compute (NCU) CSV exports.
Terminal UI for NVIDIA Nsight Systems profiles — timeline viewer, kernel navigator, NVTX hierarchy
The awesome collection of OpenClaw skills. 5,400+ skills filtered and categorized from the official OpenClaw Skills Registry.🦞
Nsight Python is a Python kernel profiling interface based on NVIDIA Nsight Tools
Framework to reduce autotune overhead to zero for well known deployments.
Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
incubator repo for CUDA-TileIR backend
Accelerating MoE with IO and Tile-aware Optimizations
A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.