Skip to content
View chenriwei's full-sized avatar
🎯
Focusing
🎯
Focusing
  • Bytedance
  • Beijing

Block or report chenriwei

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

[ACL 2024] MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues

154 117 Updated Jul 24, 2024

A framework for the evaluation of autoregressive code generation language models.

Python 1,054 261 Updated Jul 22, 2025

🐢 Open-Source Evaluation & Testing library for LLM Agents

Python 5,748 517 Updated Aug 12, 2026

Evaluation and Tracking for LLM Experiments and AI Agents

Python 3,504 319 Updated Aug 11, 2026

Supercharge Your LLM Application Evaluations 🚀

Python 15,284 1,623 Updated Feb 24, 2026

AI Observability & Evaluation

Python 11,002 1,053 Updated Aug 12, 2026

[ACL 2024 Demo] Official GitHub repo for UltraEval: An open source framework for evaluating foundation models.

Python 257 23 Updated Oct 30, 2024

🤗 Evaluate: A library for easily evaluating machine learning models and datasets.

Python 2,476 332 Updated Jul 6, 2026

The LLM Evaluation Framework

Python 17,541 1,790 Updated Aug 12, 2026

A framework for few-shot evaluation of language models.

Python 13,599 3,477 Updated Aug 11, 2026

Gorilla: Training and Evaluating LLMs for Function Calls (Tool Calls)

Python 12,991 1,398 Updated Apr 13, 2026

Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends

Python 2,516 526 Updated Aug 11, 2026

A programming framework for agentic AI

Python 60,377 9,095 Updated Apr 15, 2026

Summarize existing representative LLMs text datasets.

1,482 148 Updated Mar 11, 2026

A simple service that integrates vLLM with Ray Serve for fast and scalable LLM serving.

Python 79 13 Updated Apr 6, 2024

🏄 Scalable embedding, reasoning, ranking for images and sentences with CLIP

Python 12,834 2,067 Updated Jan 23, 2024

A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.

Python 3,223 443 Updated Aug 12, 2026

Code and implementations for the ACL 2025 paper "AgentGym: Evolving Large Language Model-based Agents across Diverse Environments" by Zhiheng Xi et al.

Python 828 115 Updated May 30, 2026

Awesome-LLM-Eval: a curated list of tools, datasets/benchmark, demos, leaderboard, papers, docs and models, mainly for Evaluation on LLMs. 一个由工具、基准/数据、演示、排行榜和大模型等组成的精选列表,主要面向基础大模型评测,旨在探求生成式AI的技术边界.

654 83 Updated Nov 24, 2025

Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷

Python 6,867 404 Updated Aug 12, 2026

Cleanlab's open-source library is the standard data-centric AI package for data quality and machine learning with messy, real-world data and labels.

Python 11,622 912 Updated Jan 13, 2026

A Lightweight Face Recognition and Facial Attribute Analysis (Age, Gender, Emotion and Race) Library for Python

Python 23,268 3,165 Updated Aug 9, 2026

text and image to video generation: CogVideoX (2024) and CogVideo (ICLR 2023)

Python 12,948 1,325 Updated Nov 4, 2025

MLRun is an open source MLOps platform for quickly building and managing continuous ML applications across their lifecycle. MLRun integrates into your development and CI/CD environment and automate…

Python 1,690 314 Updated Aug 11, 2026

😎 A curated list of awesome MLOps tools

Python 5,234 764 Updated Apr 29, 2026

📚 Papers & tech blogs by companies sharing their work on data science & machine learning in production.

30,013 3,985 Updated Jul 18, 2024

A curated list of awesome open source libraries to deploy, monitor, version and scale your machine learning

20,840 2,591 Updated Aug 10, 2026

Learn how to design systems at scale and prepare for system design interviews

45,458 5,944 Updated Jul 8, 2026

The Open Context Layer for Data and AI , OpenMetadata is the open platform for building trusted data context and business semantics for humans, AI assistants, and agents.

TypeScript 14,854 2,309 Updated Aug 12, 2026

AI Infra / AI Orchestration / AI Control Plane

MDX 3,718 330 Updated Aug 11, 2026
Next