Stars
Open-source inference server and production cluster for all the models your agent needs.
A 10-week, 30-minutes-a-day roadmap for LLM inference serving and optimization. vLLM, SGLang, quantization, speculative decoding, benchmarking.
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
The all-in-one, open-source backend platform for agentic coding. InsForge gives your coding agent database, auth, storage, compute, hosting, and AI gateway to ship full-stack apps end-to-end.
In-depth tutorials on LLMs, RAGs and real-world AI agent applications.
Agent Reinforcement Trainer: train multi-step agents for real-world tasks using GRPO. Give your agents on-the-job training. Reinforcement learning for Qwen3.6, GPT-OSS, Llama, and more!
OCR model that handles complex tables, forms, handwriting with full layout.
Sim is the collaborative workspace to build, deploy, and monitor AI agents and workflows. Used by 100,000+ builders.
Unified multimodal backend for AI data apps
An awesome & curated list of best LLMOps tools for developers
Easily fine-tune, evaluate and deploy Qwen, Gemma, or any open weight LLM!
A Kubernetes deployable instance of GroundX for document parsing, storage, and search.
Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.
The LLM's practical guide: From the fundamentals to deploying advanced LLM and RAG apps to AWS using LLMOps best practices
Command-line program to download videos from YouTube.com and other video sites
50+ tutorials and implementations for Generative AI Agent techniques, from basic conversational bots to complex multi-agent systems.
SwarmZero's SDK for building AI agents, swarms of agents and much more.
Stop renting your intelligence. Own it with AnythingLLM. Everything you need for a powerful local-first agent experience
This repository showcases various advanced techniques for Retrieval-Augmented Generation (RAG) systems. Each technique has a detailed notebook tutorial.
A minimal Python framework for building custom AI inference servers with full control over logic, batching, and scaling.
20+ high-performance LLMs with recipes to pretrain, finetune and deploy at scale.
OCR, layout analysis, reading order, table recognition in 90+ languages
A proxy server for multiple ollama instances with Key security