Stars
Distributed transactional key-value database, originally created to complement TiDB
TokenSpeed is a speed-of-light LLM inference engine.
Achieve state of the art inference performance with modern accelerators on Kubernetes
A pure-Python implementation of the Nvidia CuTe layout algebra intended to be approachable and easy to learn.
Tantivy is a full-text search engine library inspired by Apache Lucene and written in Rust
The agent that grows with you
Lightweight coding agent that runs in your terminal
Rust bindings for AppKit (macOS) and UIKit (iOS/tvOS). Experimental, but working!
For developers, who are building real-time data-driven applications, Redis is the preferred, fastest, and most feature-rich cache, data structure server, and document and vector query engine.
DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm
cuda-oxide is an experimental Rust-to-CUDA compiler that lets you write (SIMT) GPU kernels in safe(ish), idiomatic Rust. It compiles standard Rust code directly to PTX — no DSLs, no foreign languag…
A curriculum for learning about gpu performance engineering, from scratch to what the frontier AI labs do
Qwen3-omni is a natively end-to-end, omni-modal LLM developed by the Qwen team at Alibaba Cloud, capable of understanding text, audio, images, and video, as well as generating speech in real time.
SGLang-Omni empowers high-performance serving for TTS, ASR, speech and omni models.
A Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone
AI Infra / AI Orchestration / AI Control Plane
AI agent toolkit: unified LLM API, agent loop, TUI, coding agent CLI
An Interactive 3D Visual Tree Map for your S3 Bucket Storage
Experimental web-based simulator for exploring metastable behaviors in distributed systems
DeepTutor: Lifelong Personalized Tutoring. https://deeptutor.info/.
Get your documents ready for gen AI