Stars
Comprehensive AI Model Evaluation Framework with advanced techniques including Temperature-Controlled Verdict Aggregation via Generalized Power Mean. Support for multiple LLM providers and 15+ eval…
A minimal, ready-to-use scaffold for launching CrewAI projects using docker-compose. Includes basic setup, configuration, and best practices to help you hit the ground running.
Полный курс по Harness 2026 на русском языке. Все что нужно знать в одщном. метсе
Heisenbug 2026: сравнительное исследование 4 подходов к управлению контекстом AI-агентов
Automated LLM testing pipeline for LM Studio using Eval AI Library. Features dynamic model loading/unloading, interactive CLI, multiple metrics (RAG, Security, Deterministic), and integrated web da…
Evaluation framework for agentic AI systems with CI/CD-enforced quality gates, hallucination detection, and tool-call validation
WebSocket testing toolkit — CLI + Web Dashboard. Connect, record & replay, load test, validate schemas.