-
HKUST(GZ)
- Guangzhou
Starred repositories
💫 Toolkit to help you get started with Spec-Driven Development
Wrap Antigravity, ChatGPT Codex, Claude Code, Grok Build as an OpenAI/Gemini/Claude/Codex compatible API service, allowing you to enjoy the free Gemini 3.1 Pro, GPT 5.6 Series, Grok 4.5, Claude mod…
A benchmark of real-world DL kernel problems
End-to-end benchmark for AI-generated GPU kernels, drawn from real production traces — turn a PyTorch reference into a DSL kernel (Triton, Gluon, FlyDSL, CuteDSL) and grade it on compilation, numer…
An end-to-end agent project for GPU kernel implementation, analysis, profiling, and iterative optimization. It helps an agent turn PyTorch logic or an existing kernel into a high-performance GPU ke…
Mirror of https://gitcode.com/Ascend/AscendNPU-IR
Automated High-Performance GPU Kernel Generation
KernelBenchX: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels
Manage multiple AI terminal agents like Claude Code, Codex, OpenCode, and Amp.
A kernel library written in tilelang
An Evidence-Graded Catalog of Benchmarks for LLM Kernel Agents
Run Claude Code across multiple Claude accounts — a transparent proxy that auto-switches on quota and minimizes prompt-cache rebuilds
Reference code for the Meta-Harness paper.
CUDA kernels for linear attention variants, written in CuTe DSL and CUTLASS C++.
Surgical GPU kernel benchmark: 7 hard problems, frontier coding agents, roofline-graded against hardware peak.
A Claims-Annotated Catalog of LLM-Driven Kernel Generation Systems
Mobile and Web client for Codex and Claude Code, with realtime voice, encryption and fully featured
SkillOpt is a text-space optimizer that trains reusable natural-language skills for frozen LLM agents through trajectory-driven edits, validation-gated updates, and deployable best_skill.md artifacts.
Multi-account Claude proxy with automatic quota-based rotation
Scalable, Portable and Distributed Gradient Boosting (GBDT, GBRT or GBM) Library, for Python, R, Java, Scala, C and more. Runs on single machine, Hadoop, Spark, Dask, Flink and DataFlow
Ralph is an autonomous AI agent loop that runs repeatedly until all PRD items are complete.
Public skills collected from well-known open-source projects focused on LLM infrastructure, GPU kernels, compiler/operator development
Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels
Open source skill library for AI coding agents to write, optimize, and debug high performance compute kernels across CUDA, Triton, and quantized workloads.
Mirage Persistent Kernel: Compiling LLMs into a MegaKernel