-
Imperial College London
- London, UK
- http://bahri.io/
- @MehdiBahri1
Lists (1)
Sort Name ascending (A-Z)
Stars
Manages Unified Access to Generative AI Services built on Envoy Gateway
TurboQuant: Near-optimal KV cache quantization for LLM inference (3-bit keys, 2-bit values) with Triton kernels + vLLM integration
UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g., GPU-driven)
Scalable data pre processing and curation toolkit for LLMs
Evaluate and improve models and agents using environments
NeMo Guardrails is an open-source toolkit for easily adding programmable guardrails to LLM-based conversational systems.
Scalable toolkit for efficient model reinforcement
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
A Datacenter Scale Distributed Inference Serving Framework
Open source re-implementation of Single Shot End-to-end (Road) Graph Extraction (SERGE) presented at CVPR EARTHVISION 2022
A powerful MCP toolkit for coding, providing semantic retrieval and editing capabilities - the IDE for your agent
Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.
Plano is an AI-native proxy server and data plane for agentic apps. Smart LLM routing, observability, agent orchestration, and guardrails so you stay focused on your agents core logic.
📦 eCAL - enhanced Communication Abstraction Layer. A high performance publish-subscribe, client-server cross-plattform middleware.
Learn how to design large-scale systems. Prep for the system design interview. Includes Anki flashcards.
LMDeploy is a toolkit for compressing, deploying, and serving LLMs.
SGLang is a high-performance serving framework for large language models and multimodal models.
A high-throughput and memory-efficient inference and serving engine for LLMs
Triton Model Navigator is an inference toolkit designed for optimizing and deploying Deep Learning models with a focus on NVIDIA GPUs.
Python Fire is a library for automatically generating command line interfaces (CLIs) from absolutely any Python object.
CMP314 Optimizing NLP models with Amazon EC2 Inf1 instances in Amazon Sagemaker
The Triton backend for the ONNX Runtime.
Course to get into Large Language Models (LLMs) with roadmaps and Colab notebooks.
Continuously Masked Transformer for Image Inpainting, ICCV, 2023