Skip to content
View xiaoyewww's full-sized avatar

Block or report xiaoyewww

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.

Python 1,977 356 Updated Aug 13, 2026

High Performance LLM Inference Operator Library

C++ 1,107 131 Updated Aug 6, 2026

LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar

Python 18,642 1,610 Updated Aug 13, 2026

Skills for writing tilelang and debugging with CUDA toolkits.

Python 133 5 Updated May 20, 2026

Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels

Python 7,205 691 Updated Aug 12, 2026

🎨 The open-source Claude Design alternative. 🖥️ Local-first desktop app. 🖼️ Your coding agent becomes the design engine: prototypes, landing pages, dashboards, slides, images & video — real files, …

TypeScript 85,374 10,009 Updated Aug 13, 2026

Assign issues to Claude Code, Codex, Cursor, and 17 more coding agents like teammates — open-source and self-hostable.

Go 45,669 5,799 Updated Aug 12, 2026

Persistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with C…

JavaScript 90,566 7,900 Updated Aug 13, 2026

A single CLAUDE.md file to improve Claude Code behavior, derived from Andrej Karpathy's observations on LLM coding pitfalls.

201,935 20,722 Updated Apr 20, 2026

CUDA kernels for linear attention variants, written in CuTe DSL and CUTLASS C++.

Python 536 69 Updated Aug 12, 2026

期货自动交易

C 8,405 1,933 Updated Feb 28, 2026

AI agents running research on single-GPU nanochat training automatically

Python 93,752 13,296 Updated Mar 26, 2026

Tiny, Fast, and Deployable anywhere — automate the mundane, unleash your creativity

Go 29,855 4,446 Updated Aug 7, 2026

Skill + Plugin Registry for OpenClaw

TypeScript 9,296 1,446 Updated Aug 13, 2026

Byted PyTorch Distributed for Hyperscale Training of LLMs and RLs

Python 1,037 64 Updated Mar 3, 2026

Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing,…

Rust 458 136 Updated Aug 13, 2026

Fast inference from large lauguage models via speculative decoding

Python 924 95 Updated Aug 22, 2024

Machine Learning Engineering Open Book

Python 18,603 1,199 Updated Aug 13, 2026

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

JavaScript 239,777 36,392 Updated Aug 12, 2026

Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks

Python 4,335 743 Updated Aug 11, 2026

TurboDiffusion: 100–200× Acceleration for Video Diffusion Models

Python 3,607 275 Updated Aug 5, 2026

[ECCV 2026 Oral] Implementation of "Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length"

Python 2,358 277 Updated Jul 26, 2026

"AI-Trader: 100% Fully-Automated Agent-Native Trading"

Python 21,308 3,257 Updated Jun 11, 2026

linux运维相关脚本

Shell 1 Updated Sep 24, 2025

The official repo of Pai-Megatron-Patch for LLM & VLM large scale training developed by Alibaba Cloud.

Python 1,589 232 Updated Dec 15, 2025

compiler learning resources collect.

Python 2,766 369 Updated May 20, 2026

"DeepCode: Open Agentic Coding (Paper2Code & Text2Web & Text2Backend)"

Python 16,341 2,136 Updated Aug 12, 2026

Efficient implementation of DeepSeek Ops (Blockwise FP8 GEMM, MoE, and MLA) for AMD Instinct MI300X

C++ 80 7 Updated Feb 11, 2026
Next