☁️ 🚀 📊 📈 Evaluating state of the art in AI
-
Updated
Sep 20, 2026 - Python
☁️ 🚀 📊 📈 Evaluating state of the art in AI
An automatic evaluator for instruction-following language models. Human-validated, high-quality, cheap, and fast.
CogDL: A Comprehensive Library for Graph Deep Learning (WWW 2023)
ClawProBench is a live-first benchmark harness for evaluating LLM agents in the OpenClaw runtime with deterministic grading and repeated-trial reliability.
Level-Up is a Laravel package introducing gamification into your applications. Users earn experience points (XP) and levels through interactions, while also unlocking achievements. It promotes engagement, competition, and fun through its dynamic leaderboard feature. Customisable to fit your specific needs
Awesome-LLM-Eval: a curated list of tools, datasets/benchmark, demos, leaderboard, papers, docs and models, mainly for Evaluation on LLMs. 一个由工具、基准/数据、演示、排行榜和大模型等组成的精选列表,主要面向基础大模型评测,旨在探求生成式AI的技术边界.
Visually Explore the Stanford Question Answering Dataset
SKAB - Skoltech Anomaly Benchmark. Time-series data for evaluating Anomaly Detection algorithms.
A curated list of awesome leaderboard-oriented resources for AI domain
The ICPC Series Competition Leaderboard Visualization Engine
Official Problem Sets / Reference Kernels for the GPU MODE Leaderboard!
A joint community effort to create one central leaderboard for LLMs.
A large-scale (194k), Multiple-Choice Question Answering (MCQA) dataset designed to address realworld medical entrance exam questions.
OpenXAI : Towards a Transparent Evaluation of Model Explanations
Hallucinations (Confabulations) Document-Based Benchmark for RAG. Includes human-verified questions and answers.
Discover the best developers — and become one. Drop a GitHub handle for a 0–100 value & trust score in 30s: see your gaps, discover top devs, get found. Exposes PR farmers, AI bots & fork-hoarders. Deterministic scoring, self-hostable.
A leaderboard of the top open-source e-commerce platforms. Promoting the bests for building reliable stores.
The robust European language model benchmark.
AI 大模型世界,把 556 个大模型拟人化成像素小人的可视化站点。进来就能看到此刻谁最聪明、谁最会写代码、谁最便宜、谁刚发布,往下是国内与国外分区的厂商广场、完整的发布时间线和多维排行榜。搜索认模型名、厂商和能力,输入「多模态」会直接列出全部多模态模型。数据取自 Epoch AI、models.dev、LiveBench 与 Hugging Face,每小时自动同步,所有文案由真实数据生成,不调用任何 LLM。Next.js 静态导出,零后端。
powerful Discord bot that includes XP system, Leaderboard, Music, Welcome and farewell message, Moderation, and much more!
To associate your repository with the leaderboard topic, visit your repo's landing page and select "manage topics."