Stars
Autonomous GPU Kernel Generation & Optimization via Deep Agents
Code for VLM4Bio, a benchmark dataset of scientific question-answer pairs used to evaluate pretrained VLMs for trait discovery from biological images.
Official implementation of AsymFlow, pi-Flow, GMFlow
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
Decentralized LLMs fine-tuning and inference with offloading
CUDA Templates and Python DSLs for High-Performance Linear Algebra
Open-source implementation of AlphaEvolve
Achieve state of the art inference performance with modern accelerators on Kubernetes
New ML model created with Scikit Learn to recognize logical fallacies in arguments through text.
AgentFlow: In-the-Flow Agentic System Optimization
WaferLLM: Large Language Model Inference at Wafer Scale
AIInfra(AI 基础设施)指AI系统从底层芯片等硬件,到上层软件栈支持AI大模型训练和推理。
AG2 (formerly AutoGen): The Open-Source AgentOS.Join us at: https://discord.gg/sNGSwQME3x
This project develops a high-performance KV-cache management framework for multi-document RAG tasks. It focuses on reducing time-per-output-token (TPOT) and improving throughput through adaptive ca…
Distributed Inference for Large Language Models.
verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
aider is AI pair programming in your terminal
[ICLR 2025🔥] SVD-LLM & [NAACL 2025🔥] SVD-LLM V2