Skip to content
View googs1025's full-sized avatar
🎯
🎯

Organizations

@kubernetes @kubernetes-sigs @volcano-sh @koordinator-sh @InftyAI @llm-d

Block or report googs1025

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Agent Skills & Tooling

Claude Skills 生态、agent harness 工具
48 repositories

AI Agent Frameworks

Claude/LangChain/LangGraph/MCP/Agent SDK
113 repositories

AI Infra Learning (中文)

中文 AI infra/LLM 教程与 awesome list
33 repositories

Container Runtimes / Wasm

containerd/runc/kata/youki/wasm
27 repositories

CS Fundamentals / Interview

算法、系统设计、面试题、经典书
29 repositories

Go & Rust Foundations

Go/Rust 语言学习、基础库、Web 框架
89 repositories

K8s × AI Serving

K8s 上跑 LLM/AI 工作负载的平台
45 repositories

K8s Core & Controllers

K8s 主线、controllers、operator SDK、CRD、kubectl、client-go
360 repositories

Starred repositories

Showing results

A local-first routing and coordination engine for long-running agent work.

Rust 20 3 Updated Sep 24, 2026

Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding,…

Go 212 34 Updated Sep 24, 2026
TypeScript 2 1 Updated Sep 17, 2026

GPT-2-style LLM built from scratch in C/CUDA with hand-written backprop, BPE tokenizer, FlashAttention, pretraining, and SFT.

Cuda 121 23 Updated Jun 18, 2026

Generate text, images, video, speech, and music by MiniMax.

TypeScript 2,171 185 Updated Sep 19, 2026

Run frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦

C 37,401 4,047 Updated Sep 23, 2026

Orca is the ADE for working with a fleet of parallel agents. Run any coding agent with your own subscription. Available on desktop, mobile and remote runtime.

TypeScript 76,856 5,032 Updated Sep 24, 2026

Agent skill that removes signs of AI-generated writing from text

Python 51,722 4,140 Updated Sep 6, 2026
Python 158 21 Updated Sep 22, 2026

《深入理解 AI Agent:设计原理与工程实践》(李博杰 著)开源主仓库:全书正文、编译版 PDF 与按章配套代码

Python 50,614 5,681 Updated Sep 24, 2026

vLLM plugin for attention-ffn disaggregation support

Python 247 51 Updated Sep 24, 2026

《深入理解 AI Infra:量化分析与系统设计》(李博杰 著)开源书稿:从硬件约束和模型架构出发,量化推导 LLM 推理与训练系统设计。含全书正文、PDF、配套计算工具与实验

Python 5,185 378 Updated Sep 23, 2026

An agent harness that compiles a model into one provably-correct, self-retargeting CUDA megakernel and self-tunes it past cuBLAS at batch-1 LLM decode, paper: https://arxiv.org/abs/2606.09682

Python 142 15 Updated Sep 18, 2026

A transparent, in-container GPU resource controller that enforces memory and compute limits by intercepting CUDA calls without application or driver changes.

C 336 214 Updated Sep 24, 2026

SGLang-Omni is a high-performance serving framework for audio models (TTS, ASR) and unified multimodal models.

Python 1,268 521 Updated Sep 24, 2026

Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end.

Python 3,420 412 Updated Sep 23, 2026

CUDA Library Samples

Cuda 2,516 478 Updated Sep 23, 2026

Optimized primitives for collective multi-GPU communication

C++ 5,114 1,426 Updated Sep 23, 2026

Ecosystem of libraries and tools for writing and executing fast GPU code fully in Rust.

Rust 5,394 250 Updated Sep 22, 2026

A modern replacement for Redis and Memcached

C++ 31,669 1,268 Updated Sep 24, 2026

Nephio is a Kubernetes-based automation platform for deploying and managing highly distributed, interconnected workloads such as 5G Network Functions, and the underlying infrastructure on which tho…

Go 182 79 Updated Sep 21, 2026
Python 6 16 Updated Aug 13, 2026

LLM inference in C/C++

C++ 129,373 23,666 Updated Sep 24, 2026

Tensors and Dynamic neural networks in Python with strong GPU acceleration

Python 103,230 30,129 Updated Sep 24, 2026

Go bindings to systemd socket activation, journal, D-Bus, and unit files

Go 2,715 339 Updated Jul 23, 2026

Kubernetes controllers for fast model actuation using vLLM sleep/wake and launcher-based model swapping

Go 25 24 Updated Sep 24, 2026

Make Every Team AI Native

TypeScript 4,984 371 Updated Sep 24, 2026

学摩尔线程 MUSA SDK 的学习记录

mupad 3 1 Updated Sep 22, 2026

how to optimize some algorithm in cuda.

Cuda 3,289 292 Updated Sep 14, 2026
Next