Stars
A Child-Safety Risk Benchmark for Language Models
An Analysis of Active Learning Algorithms using Real-World Crowd-sourced Text Annotations
AI Benchmark for Investment Banking Workflows
Packaging architecture for RL environments
Agent-as-a-Judge grading framework for evaluating AI outputs/deliverables
Post-training framework for large models, from new objectives to new rollout systems.
Score the trustworthiness of outputs from any LLM in real-time
A Structured Output Benchmark whose 'ground-truth' is actually right
A self-serve demo and walkthrough of the Cleanlab AI Platform
Build data processing and data analysis pipelines that leverage the power of LLMs 🧠
The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while control…
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command li…
Inference-time scaling for LLMs-as-a-judge.
[JMLR 2026] "UQLM: A Python Package for Uncertainty Quantification in Large Language Models"
Langtrace 🔍 is an open-source, Open Telemetry based end-to-end observability tool for LLM applications, providing real-time tracing, evaluations and metrics for popular LLMs, LLM frameworks, vector…
In-depth tutorials on LLMs, RAGs and real-world AI agent applications.
Generating a trustworthiness or reliability score of a large language model's response for both direct questions and retrieval augment generation (RAG)
Python client library for Cleanlab Trustworthy Language Model
Python client to integrate Cleanlab Codex with your AI Agent
Tutorial on Automated Machine Learning at KDD 2020
Bandit algorithms for dynamic pricing of many products
List of papers on hallucination detection in LLMs.
Jupyter Notebooks to help you get hands-on with Pinecone vector databases
A Community-Driven Mapping of AI Development Tools