Skip to content
View james0zan's full-sized avatar

Organizations

@kvcache-ai

Block or report james0zan

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Embodied AI Operating System (EAIOS)

Rust 335 50 Updated Aug 19, 2026

AgentENV (AENV) is a distributed platform for running agent environments at scale.

Rust 3,242 276 Updated Aug 20, 2026

SGLang is a high-performance serving framework for large language models and multimodal models.

Python 32,138 8,035 Updated Aug 20, 2026

A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations

Python 19,257 1,530 Updated Aug 19, 2026

A PyTorch native library for training speculative decoding models

Python 227 60 Updated Aug 19, 2026

Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing,…

Rust 467 139 Updated Aug 20, 2026

Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.

Python 2,173 381 Updated Aug 20, 2026

A visualized theorem prover based on Lean 4

TypeScript 8 Updated Nov 12, 2025

Proof the completeness of Russel's Axiomatic System in lean4, and using C++ to automatically convert lean4 file to markdown file

Lean 4 Updated Jan 6, 2026
Go 98 9 Updated Sep 15, 2025

Checkpoint-engine is a simple middleware to update model weights in LLM inference engines

Python 999 102 Updated Aug 12, 2026

A Distributed Attention Towards Linear Scalability for Ultra-Long Context, Heterogeneous Data Training

Python 918 66 Updated Aug 19, 2026

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python 11,225 1,733 Updated Aug 20, 2026

I created a claude code deep researcher that seems to work better than the current deep research models

Python 147 26 Updated Jun 9, 2026

MoBA: Mixture of Block Attention for Long-Context LLMs

Python 2,166 158 Updated Apr 3, 2025

High-speed Large Language Model Serving for Local Deployment

C++ 9,722 592 Updated May 11, 2026

Merico Build is a web app empowering open source developers, maintainers, and communities with metrics from Git, GitHub, and more.

494 24 Updated Jun 29, 2021

CSI driver to bring SPDK to Kubernetes storage through NVMe-oF or iSCSI. Supports dynamic volume provisioning and enables Pods to use SPDK storage transparently.

Go 88 45 Updated Jan 27, 2026

A RocksDB compatible KV storage engine with better performance

C++ 2,152 213 Updated Jul 13, 2026

Concurrent data structures in C++

C++ 1,459 160 Updated May 16, 2026

Bot Framework provides the most comprehensive experience for building conversation applications.

JavaScript 7,807 2,421 Updated Dec 29, 2025

Multithreaded HTTP Download Accelerator

C 23 12 Updated Jul 27, 2014

A Python library for using the duoshuo API

Python 3 Updated Jul 22, 2012

A Python library for using the duoshuo API

Python 88 31 Updated Nov 23, 2021

PyCoder's Weekly Chinese Translate Sources Repo

HTML 394 92 Updated Dec 3, 2017