Skip to content
View yyccli's full-sized avatar

Block or report yyccli

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

CLI proxy that reduces LLM token consumption by 60-90% on common dev commands. Single Rust binary, zero dependencies

Rust 76,247 4,795 Updated Aug 15, 2026

Lightweight coding agent that runs in your terminal

Rust 106,142 16,115 Updated Aug 16, 2026

The open source coding agent.

TypeScript 197,829 25,481 Updated Aug 16, 2026

Tile-Based Runtime for Ultra-Low-Latency LLM Inference

Python 1,697 120 Updated Aug 13, 2026

FlyDSL is the Python front‑end of the project: a Flexible Layout Python DSL for expressing tiling, partitioning, data movement, and kernel structure at a high level.

Python 261 108 Updated Aug 16, 2026

An agentic skills framework & software development methodology that works.

Shell 272,527 24,370 Updated Aug 13, 2026

high-performance linear attention kernel library built on TileLang

Python 638 66 Updated Aug 11, 2026

KernelBench: Can LLMs Write GPU Kernels? - Benchmark + Toolkit with Torch -> CUDA (+ more DSLs)

Jupyter Notebook 1,199 189 Updated Mar 24, 2026

FlashKDA: high-performance Kimi Delta Attention kernels

Cuda 1,213 116 Updated Jul 30, 2026

High-performance GEMM kernel examples with FlyDSL on AMD GPUs.

Python 28 2 Updated Aug 13, 2026

A cross-platform desktop All-in-One assistant for Claude Code, Codex, OpenCode, OpenClaw, Grok Build & Hermes Agent. Only official website: ccswitch.io

Rust 127,432 8,700 Updated Aug 16, 2026

Claude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining complex code, and handling git workflo…

Python 141,579 22,730 Updated Aug 14, 2026

Autonomous GPU Kernel Generation & Optimization via Deep Agents

Python 512 86 Updated Jul 15, 2026

CUDA kernels for linear attention variants, written in CuTe DSL and CUTLASS C++.

Python 538 69 Updated Aug 14, 2026

🚀 Efficient implementations for emerging model architectures

Python 5,561 659 Updated Aug 14, 2026

Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞

TypeScript 386,416 81,213 Updated Aug 16, 2026
Python 156 18 Updated Aug 8, 2026
Python 21 2 Updated Mar 17, 2026

[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.

Cuda 3,645 480 Updated Jan 17, 2026

A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.

Python 4,747 787 Updated May 17, 2026

RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.

Cuda 1,304 258 Updated Aug 15, 2026

Mirage Persistent Kernel: Compiling LLMs into a MegaKernel

Cuda 2,425 238 Updated Aug 4, 2026

CUDA Python: Performance meets Productivity

Cython 3,344 321 Updated Aug 16, 2026

An Emacs framework for the stubborn martian hacker

Emacs Lisp 22,584 3,158 Updated Aug 6, 2026

DeepEP: an efficient expert-parallel communication library

Cuda 9,993 1,380 Updated Aug 5, 2026

NVIDIA Inference Xfer Library (NIXL)

C++ 1,193 404 Updated Aug 15, 2026

A Datacenter Scale Distributed Inference Serving Framework

Rust 7,773 1,443 Updated Aug 16, 2026

AI Tensor Engine for ROCm

Python 528 481 Updated Aug 16, 2026

[DEPRECATED] Moved to ROCm/rocm-libraries repo. NOTE: develop branch is maintained as a read-only mirror

C++ 543 304 Updated Aug 15, 2026

DeepGEMM: clean and efficient BLAS kernel library on GPU

Cuda 7,683 1,172 Updated Aug 11, 2026
Next