Skip to content
View acelyc111's full-sized avatar
:octocat:
working
:octocat:
working

Organizations

@apache @XiaoMi @pegasus-kv

Block or report acelyc111

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

A rule-based tunnel for Android.

Kotlin 44,285 2,787 Updated Aug 8, 2026

TokenSpeed is a speed-of-light LLM inference engine.

Python 1,836 225 Updated Aug 9, 2026

便宜机场,一元机场,性价比机场,白嫖机场,免费机场,机场推荐,github加速,github文件加速,机场订阅加速,2025年最新科学上网,vpn机场推荐

699 14 Updated Aug 5, 2026

Code, labs, and resources for O'Reilly AI Systems Performance Engineering: GPU optimization, distributed training, inference scaling, and full-stack tuning.

Python 1,785 251 Updated Aug 4, 2026

CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation

Python 1,128 95 Updated Jul 8, 2026

Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞

TypeScript 385,639 81,058 Updated Aug 9, 2026

The open source coding agent.

TypeScript 195,252 25,010 Updated Aug 9, 2026

Alibaba Cloud's high-performance KVCache system for LLM inference, with components for global cache management, inference simulation(HiSim), and more.

C++ 225 51 Updated Aug 7, 2026

🌟100+ 原创 LLM / RL 原理图📚,《大模型算法》作者巨献!💥(100+ LLM/RL Algorithm Maps )

Python 4,752 455 Updated Jul 27, 2026

mimalloc is a compact general purpose allocator with excellent performance.

C 13,269 1,151 Updated Aug 9, 2026

One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks

Python 4,353 636 Updated Aug 6, 2026

High Performance LLM Inference Operator Library

C++ 1,094 127 Updated Aug 6, 2026

[CVPR 2025] A Comprehensive Benchmark for Document Parsing and Evaluation

Python 1,958 191 Updated Jul 27, 2026

A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.

Python 4,715 778 Updated May 17, 2026

[NeurIPS 2025 D&B] 🚀 SWE-bench Goes Live!

Python 217 31 Updated Jun 11, 2026

Fast Hadamard transform in CUDA, with a PyTorch interface

C 343 69 Updated Mar 10, 2026

🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.

Python 34,269 7,228 Updated Aug 9, 2026

how to optimize some algorithm in cuda.

Cuda 3,189 289 Updated Aug 7, 2026
Python 233 14 Updated Sep 25, 2025

DeepXTrace is a lightweight tool for precisely diagnosing slow ranks in DeepEP-based environments.

Python 102 7 Updated Jan 16, 2026

A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.

C++ 1,513 276 Updated Aug 7, 2026

🔊 让小爱音箱「听见你的声音」,解锁无限可能。

Rust 2,594 377 Updated Apr 4, 2026

🏠 将小爱音箱接入 ChatGPT 和豆包,改造成你的专属语音助手。

TypeScript 12,512 1,757 Updated Apr 4, 2026

AIInfra(AI 基础设施)指AI系统从底层芯片等硬件,到上层软件栈支持AI大模型训练和推理。

Jupyter Notebook 7,885 1,014 Updated Dec 22, 2025

Materials for learning SGLang

869 68 Updated Jan 5, 2026

A Datacenter Scale Distributed Inference Serving Framework

Rust 7,713 1,421 Updated Aug 9, 2026

🚀 Efficient implementations for emerging model architectures

Python 5,528 642 Updated Aug 9, 2026

Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels

Python 7,165 682 Updated Aug 9, 2026

LLM inference in C/C++

C++ 123,168 21,476 Updated Aug 9, 2026

Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.

C++ 6,208 1,066 Updated Aug 9, 2026
Next