Skip to content
View jwmueller's full-sized avatar

Organizations

@cleanlab

Block or report jwmueller

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

A Child-Safety Risk Benchmark for Language Models

Python 3 Updated Jun 30, 2026

An Analysis of Active Learning Algorithms using Real-World Crowd-sourced Text Annotations

3 Updated Apr 5, 2026

Website to host ACM CAIS EL EVAL Workshop

HTML 2 1 Updated May 26, 2026

AI Benchmark for Investment Banking Workflows

Python 40 2 Updated Jun 30, 2026

Packaging architecture for RL environments

Dockerfile 5 Updated Apr 16, 2026

Agent-as-a-Judge grading framework for evaluating AI outputs/deliverables

Python 51 6 Updated Aug 7, 2026

Post-training framework for large models, from new objectives to new rollout systems.

Python 213 18 Updated Aug 5, 2026

Score the trustworthiness of outputs from any LLM in real-time

Python 7 4 Updated Jul 6, 2026

A Structured Output Benchmark whose 'ground-truth' is actually right

Jupyter Notebook 25 1 Updated Jul 24, 2026

A self-serve demo and walkthrough of the Cleanlab AI Platform

TypeScript 2 Updated Oct 29, 2025

Build data processing and data analysis pipelines that leverage the power of LLMs 🧠

Python 261 8 Updated Jul 26, 2026

The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while control…

Python 27,424 6,128 Updated Aug 8, 2026

🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23

TypeScript 32,756 3,517 Updated Aug 9, 2026

Repository for the TLM core algorithm

Python 1 1 Updated Dec 2, 2025

Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command li…

TypeScript 24,074 2,167 Updated Aug 8, 2026

Inference-time scaling for LLMs-as-a-judge.

Jupyter Notebook 346 28 Updated Nov 5, 2025

[JMLR 2026] "UQLM: A Python Package for Uncertainty Quantification in Large Language Models"

Python 1,188 129 Updated Aug 3, 2026
Python 498 72 Updated Aug 4, 2026

AI Observability & Evaluation

Python 10,953 1,044 Updated Aug 8, 2026

Langtrace 🔍 is an open-source, Open Telemetry based end-to-end observability tool for LLM applications, providing real-time tracing, evaluations and metrics for popular LLMs, LLM frameworks, vector…

TypeScript 1,224 125 Updated Nov 17, 2025

In-depth tutorials on LLMs, RAGs and real-world AI agent applications.

Jupyter Notebook 36,907 6,092 Updated Jul 27, 2026

Generating a trustworthiness or reliability score of a large language model's response for both direct questions and retrieval augment generation (RAG)

Jupyter Notebook 2 Updated Feb 5, 2025

Python client library for Cleanlab Trustworthy Language Model

Python 24 2 Updated Dec 9, 2025

Python client to integrate Cleanlab Codex with your AI Agent

Python 19 4 Updated Nov 19, 2025

Tutorial on Automated Machine Learning at KDD 2020

Jupyter Notebook 55 25 Updated Sep 1, 2020
OpenEdge ABL 5 2 Updated Dec 11, 2014

Bandit algorithms for dynamic pricing of many products

Python 42 13 Updated Nov 5, 2019

List of papers on hallucination detection in LLMs.

1,123 91 Updated Jul 24, 2026

Jupyter Notebooks to help you get hands-on with Pinecone vector databases

Jupyter Notebook 3,035 1,073 Updated Aug 4, 2026

A Community-Driven Mapping of AI Development Tools

HTML 637 118 Updated Jul 24, 2026
Next