Skip to content
#

ai-benchmark

Here are 148 public repositories matching this topic...

benchmark-radar

Track 17,170+ AI benchmark, eval, dataset, and data-quality records from 37 public sources, with linked evidence and daily updates.

  • Updated Sep 22, 2026
  • Python

AI clothes swap prompt cookbook with 50 virtual try-on prompts, KIE GPT Image before/after examples, failure fixes, and a benchmark rubric for AIClothSwap.

  • Updated Jul 9, 2026
  • HTML

🤖 A curated list of resources for testing AI agents - frameworks, methodologies, benchmarks, tools, and best practices for ensuring reliable, safe, and effective autonomous AI systems

  • Updated May 28, 2025

MindTrial: Evaluate and compare AI language models (LLMs) on text-based tasks with optional file/image attachments and tool use. Supports multiple providers (OpenAI, Google, Anthropic, DeepSeek, Mistral AI, xAI, Alibaba, Moonshot AI, OpenRouter), custom tasks in YAML, and HTML/CSV/JSON reports.

  • Updated Sep 22, 2026
  • Go

Benchmark abierto en español de modelos de IA para negocios y agentes, con juez independiente (Phi-4). Calidad, costo, velocidad, contexto largo, trabajo agéntico y fuga de credenciales, cada uno por separado. Calculadora interactiva con tus propios pesos.

  • Updated Sep 21, 2026
  • Python

Add this topic to your repo

To associate your repository with the ai-benchmark topic, visit your repo's landing page and select "manage topics."

Learn more