Skip to content
View ClawSeven's full-sized avatar
👋
Hi, I am zehuan !
👋
Hi, I am zehuan !
  • AntGroup
  • Shanghai
  • 11:05 (UTC +08:00)

Block or report ClawSeven

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Starred repositories

Showing results

Executable, measurable, and reproducible AI4AI toward recursive self-improvement. Home of OpenMLE and Frontis-MA1.

Python 469 44 Updated Aug 12, 2026

AI agents running research on single-GPU nanochat training automatically

Python 93,815 13,303 Updated Mar 26, 2026

KernelBench: Can LLMs Write GPU Kernels? - Benchmark + Toolkit with Torch -> CUDA (+ more DSLs)

Jupyter Notebook 1,197 187 Updated Mar 24, 2026

LLM-as-a-Verifier is a general-purpose framework that provides fine-grained feedback for any agent without requiring additional training. It achieves SOTA performance across coding, robotics, and m…

Python 635 57 Updated Aug 7, 2026

Programmable datacenter-scale infrastructure for Agents.

Python 37 9 Updated Aug 13, 2026
C++ 116 19 Updated Aug 14, 2026

分享AI Infra知识&代码练习:PyTorch、vLLM/SGLang、slime/vime框架入门⚡️、性能加速🚀、大模型基础🧠、AI软硬件🔧等

Jupyter Notebook 3,505 331 Updated Aug 7, 2026

Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models

Python 4,589 351 Updated Jan 14, 2026

TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. Tensor…

Python 14,379 2,660 Updated Aug 14, 2026

Achieve state of the art inference performance with modern accelerators on Kubernetes

Shell 4,025 678 Updated Aug 13, 2026

An Extensible Deep Learning Library

Python 2,373 409 Updated Jul 8, 2026

[MLsys2026]: RAG on Everything with LEANN. Enjoy 97% storage savings while running a fast, accurate, and 100% private RAG application on your personal device.

Python 12,785 1,144 Updated Jul 31, 2026

Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond

Python 1,131 130 Updated Aug 13, 2026

DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm

C 21,330 1,945 Updated Aug 9, 2026

Distribute and run AI workloads on Kubernetes magically in Python, like PyTorch for ML infra.

Python 1,224 60 Updated May 29, 2026

RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.

Cuda 1,301 257 Updated Aug 14, 2026

The agent that grows with you

Python 230,200 45,563 Updated Aug 14, 2026

JAX-Toolbox

Python 426 80 Updated Aug 13, 2026

Training library for Megatron-based models with bidirectional Hugging Face conversion capability

Python 858 454 Updated Aug 14, 2026

Ongoing research training transformer models at scale

Python 17,422 4,358 Updated Aug 14, 2026

JaxPP is a library for JAX that enables flexible MPMD pipeline parallelism for large-scale LLM training

Python 83 4 Updated Aug 13, 2026

Composable transformations of Python+NumPy programs: differentiate, vectorize, JIT to GPU/TPU, and more

Python 36,157 3,736 Updated Aug 14, 2026

Tokamax: A GPU and TPU kernel library.

Python 262 47 Updated Aug 14, 2026

Showcase JaxPP with MaxText

Python 8 1 Updated Apr 2, 2026

A machine learning compiler for GPUs, CPUs, and ML accelerators

C++ 4,467 893 Updated Aug 14, 2026

Orbax provides common checkpointing and persistence utilities for JAX users

Python 528 101 Updated Aug 13, 2026

A profiling and performance analysis tool for machine learning

C++ 569 96 Updated Aug 14, 2026

TPU inference for vLLM, with unified JAX and PyTorch support.

Python 406 285 Updated Aug 14, 2026
Next