Stars
[ACL 2024] MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues
A framework for the evaluation of autoregressive code generation language models.
🐢 Open-Source Evaluation & Testing library for LLM Agents
Evaluation and Tracking for LLM Experiments and AI Agents
Supercharge Your LLM Application Evaluations 🚀
[ACL 2024 Demo] Official GitHub repo for UltraEval: An open source framework for evaluating foundation models.
🤗 Evaluate: A library for easily evaluating machine learning models and datasets.
A framework for few-shot evaluation of language models.
Gorilla: Training and Evaluating LLMs for Function Calls (Tool Calls)
Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends
Summarize existing representative LLMs text datasets.
A simple service that integrates vLLM with Ray Serve for fast and scalable LLM serving.
🏄 Scalable embedding, reasoning, ranking for images and sentences with CLIP
A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.
Code and implementations for the ACL 2025 paper "AgentGym: Evolving Large Language Model-based Agents across Diverse Environments" by Zhiheng Xi et al.
Awesome-LLM-Eval: a curated list of tools, datasets/benchmark, demos, leaderboard, papers, docs and models, mainly for Evaluation on LLMs. 一个由工具、基准/数据、演示、排行榜和大模型等组成的精选列表,主要面向基础大模型评测,旨在探求生成式AI的技术边界.
Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷
Cleanlab's open-source library is the standard data-centric AI package for data quality and machine learning with messy, real-world data and labels.
A Lightweight Face Recognition and Facial Attribute Analysis (Age, Gender, Emotion and Race) Library for Python
text and image to video generation: CogVideoX (2024) and CogVideo (ICLR 2023)
MLRun is an open source MLOps platform for quickly building and managing continuous ML applications across their lifecycle. MLRun integrates into your development and CI/CD environment and automate…
😎 A curated list of awesome MLOps tools
📚 Papers & tech blogs by companies sharing their work on data science & machine learning in production.
A curated list of awesome open source libraries to deploy, monitor, version and scale your machine learning
Learn how to design systems at scale and prepare for system design interviews
The Open Context Layer for Data and AI , OpenMetadata is the open platform for building trusted data context and business semantics for humans, AI assistants, and agents.
AI Infra / AI Orchestration / AI Control Plane