Skip to content
View HandH1998's full-sized avatar

Block or report HandH1998

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

A vLLM patch + hand‑written SM120 SASS kernels: 2‑bit MoE experts + an FP4 "delta" cache that recovers precision — matching the official (NV)FP4 checkpoint's quality on consumer Blackwell cards

Sass 517 49 Updated Aug 12, 2026

Reference implementation and examples of the CuTe Layout representation and algebra.

Python 265 25 Updated Aug 6, 2026

Orca is the ADE for working with a fleet of parallel agents. Run any coding agent with your own subscription. Available on desktop, mobile and VPS.

TypeScript 43,586 3,041 Updated Aug 12, 2026

Bridge Feishu/Lark to AI coding CLIs — Claude Code, Codex, Gemini, OpenCode… every DM, group or topic spawns its own live-streaming CLI session

TypeScript 1,052 203 Updated Aug 12, 2026

From Automated Idea Factory to Realization

Shell 1,385 124 Updated Jul 18, 2026

TurboQuant: Near-optimal KV cache quantization for LLM inference (3-bit keys, 2-bit values) with Triton kernels + vLLM integration

Python 1,720 192 Updated Mar 27, 2026

Persistent Claude/Codex terminal and Agent Workspace dashboard backed by ttyd + tmux.

Python 7 Updated Aug 9, 2026

TokenSpeed is a speed-of-light LLM inference engine.

Python 1,850 234 Updated Aug 12, 2026
C++ 5 Updated Apr 29, 2026

A kernel library written in tilelang

Python 1,717 154 Updated Apr 23, 2026

Efficient and unified implementations for TopK-based sparse attention

Cuda 39 1 Updated Jul 10, 2026
Python 89 15 Updated Apr 18, 2025

An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.

Rust 195,063 109,187 Updated Aug 6, 2026

CUDA Tile IR is an MLIR-based intermediate representation and compiler infrastructure for CUDA kernel optimization, focusing on tile-based computation patterns and optimizations targeting NVIDIA te…

C++ 1,007 85 Updated Jul 22, 2026

Bridge local AI coding agents (Claude Code, Cursor, Gemini CLI, Codex) to messaging platforms (Feishu/Lark, DingTalk, Slack, Telegram, Discord, LINE, WeChat Work). Chat with your AI dev assistant f…

Go 14,877 1,451 Updated Aug 7, 2026

PTX ISA 9.1 documentation converted to searchable markdown. Includes Claude Code skill for CUDA development.

Python 222 41 Updated Dec 24, 2025

Agentic Kernel Optimization for All — automated GPU kernel optimization for any kernel, any hardware, any language

Python 347 28 Updated May 31, 2026

Automated CUDA kernel performance diagnostics from NVIDIA Nsight Compute (NCU) CSV exports.

Rust 35 3 Updated Mar 18, 2026

Terminal UI for NVIDIA Nsight Systems profiles — timeline viewer, kernel navigator, NVTX hierarchy

Python 73 20 Updated Aug 11, 2026

The awesome collection of OpenClaw skills. 5,400+ skills filtered and categorized from the official OpenClaw Skills Registry.🦞

51,905 4,996 Updated Aug 9, 2026
TypeScript 8 Updated Mar 23, 2026

Nsight Python is a Python kernel profiling interface based on NVIDIA Nsight Tools

Python 286 21 Updated Aug 12, 2026

Framework to reduce autotune overhead to zero for well known deployments.

Python 101 16 Updated Sep 19, 2025

GPTQ inference Triton kernel

Jupyter Notebook 323 21 Updated May 18, 2023

High Performance LLM Inference Operator Library

C++ 1,107 131 Updated Aug 6, 2026

Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models

Python 4,587 350 Updated Jan 14, 2026

incubator repo for CUDA-TileIR backend

MLIR 153 15 Updated Jul 10, 2026

Accelerating MoE with IO and Tile-aware Optimizations

Python 739 95 Updated Jul 4, 2026

A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.

Python 4,738 785 Updated May 17, 2026
Next