Skip to content
@kvcache-ai

kvcache.ai

KVCache.AI is an open source orgnization between MADSys and top industry collaborators, focusing on efficient Agent/LLM serving.

Pinned Loading

  1. Mooncake Mooncake Public

    Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.

    C++ 6.1k 1k

  2. ktransformers ktransformers Public

    A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations

    Python 19.1k 1.5k

  3. AgentENV AgentENV Public

    AgentENV (AENV) is a distributed platform for running agent environments at scale.

    Rust 1.2k 101

Repositories

Showing 10 of 17 repositories
  • kvcache-blog Public
    kvcache-ai/kvcache-blog's past year of commit activity
    JavaScript 22 MIT 16 0 1 Updated Jul 28, 2026
  • Mooncake Public

    Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.

    kvcache-ai/Mooncake's past year of commit activity
    C++ 6,062 Apache-2.0 1,027 207 (1 issue needs help) 278 Updated Jul 28, 2026
  • AgentENV Public

    AgentENV (AENV) is a distributed platform for running agent environments at scale.

    kvcache-ai/AgentENV's past year of commit activity
    Rust 1,211 MIT 101 16 8 Updated Jul 28, 2026
  • sglang Public Forked from sgl-project/sglang

    SGLang is a fast serving framework for large language models and vision language models.

    kvcache-ai/sglang's past year of commit activity
    Python 14 Apache-2.0 7,499 0 16 Updated Jul 28, 2026
  • kvcache-ai/kvcache-blog-private's past year of commit activity
    JavaScript 0 MIT 15 0 0 Updated Jul 25, 2026
  • ktransformers Public

    A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations

    kvcache-ai/ktransformers's past year of commit activity
    Python 19,077 Apache-2.0 1,489 456 (1 issue needs help) 16 Updated Jul 24, 2026
  • vllm Public Forked from vllm-project/vllm

    A high-throughput and memory-efficient inference and serving engine for LLMs

    kvcache-ai/vllm's past year of commit activity
    Python 17 Apache-2.0 20,197 0 0 Updated Jul 1, 2026
  • Model-Optimizer Public Forked from NVIDIA/Model-Optimizer

    A unified library of SOTA model optimization techniques like quantization, pruning, distillation, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.

    kvcache-ai/Model-Optimizer's past year of commit activity
    Python 1 Apache-2.0 519 0 0 Updated Jun 17, 2026
  • accelerate Public Forked from huggingface/accelerate

    🚀 A simple way to launch, train, and use PyTorch models on almost any device and distributed configuration, automatic mixed precision (including fp8), and easy-to-configure FSDP and DeepSpeed support

    kvcache-ai/accelerate's past year of commit activity
    Python 2 Apache-2.0 1,453 0 1 Updated May 9, 2026
  • transformers Public Forked from huggingface/transformers

    🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.

    kvcache-ai/transformers's past year of commit activity
    Python 2 Apache-2.0 34,786 0 1 Updated May 9, 2026

Top languages

Loading…

Most used topics

Loading…