- Bay Area, CA
-
05:53
(UTC -07:00) - https://zhyncs.com
- https://orcid.org/0009-0006-7743-2508
- @zhyncs42
Stars
AI agent toolkit: unified LLM API, agent loop, TUI, coding agent CLI
AgentENV (AENV) is a distributed platform for running agent environments at scale.
MoonEP: A Perfectly Balanced Expert Parallelism Library via Dynamic Redundant Experts
Inspect: A framework for large language model evaluations
TokenSpeed is a speed-of-light LLM inference engine.
A PyTorch native library for training speculative decoding models
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing,…
A program which ensures source code files have copyright license headers by scanning directory patterns recursively
Lightweight coding agent that runs in your terminal
🐶 Kubernetes CLI To Manage Your Clusters In Style!
The easiest, most secure way to use WireGuard and 2FA.
DeepGEMM: clean and efficient BLAS kernel library on GPU
DeepEP: an efficient expert-parallel communication library
FlashMLA: Efficient Multi-head Latent Attention Kernels
Fast, Flexible and Portable Structured Generation
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. Tensor…
Official Implementation of EAGLE-1 (ICML'24), EAGLE-2 (EMNLP'24), and EAGLE-3 (NeurIPS'25).
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
Development repository for the Triton language and compiler
FlashInfer: Kernel Library for LLM Serving
CUDA Templates and Python DSLs for High-Performance Linear Algebra
Fast and memory-efficient exact attention
brpc is an Industrial-grade RPC framework using C++ Language, which is often used in high performance system such as Search, Storage, Machine learning, Advertisement, Recommendation etc. "brpc" mea…