Binary evals. Trace-centric. Error-analysis-first.
A command-line tool for evaluating LLM outputs using binary pass/fail judgments, built on Hamel Husain's evaluation principles.
π Full LLM Guide | π Quick Start | π GitHub Pages
Most teams struggle with LLM evaluation because they:
- β Use complex 1-5 scales (hard to agree on)
- β Skip manual error analysis (miss critical failures)
- β Start with expensive LLM-as-judge (waste money)
- β Build evals before understanding failures (measure wrong things)
EmbedEval fixes this with Hamel Husain's proven approach:
- β Binary only - PASS or FAIL, no debating
- β Error analysis first - Look at traces before automating
- β Cheap evals first - Assertions before LLM-as-judge
- β Trace-centric - Complete session records
- β Single annotator - "Benevolent dictator" model
Option 1: Quick Install (Recommended)
curl -fsSL https://raw.githubusercontent.com/Algiras/embedeval/main/install.sh | bashOption 2: npm (Global)
npm install -g embedevalOption 3: npx (No Install)
npx embedeval <command># 1. COLLECT - Import your LLM traces
embedeval collect ./production-logs.jsonl --output traces.jsonl
# 2. ANNOTATE - Manual error analysis (30 min for 50-100 traces)
embedeval annotate traces.jsonl --user "expert@company.com"
# 3. TAXONOMY - Build failure taxonomy
embedeval taxonomy build --annotations annotations.jsonlThat's it. You'll now see:
Pass Rate: 73%
Top Failure Categories:
1. Hallucination: 12 traces (44%)
2. Incomplete: 8 traces (30%)
3. Wrong Format: 5 traces (19%)
Watch a complete 5-minute tutorial covering the entire EmbedEval workflow:
# View the tutorial video
open embedeval-promo/out/embedeval-tutorial-full-v3.mp4What's Covered:
- β Installation verification
- β First trace collection
- β Interactive annotation (keyboard shortcuts)
- β Building failure taxonomy
- β Running evaluations
- β Understanding Hamel Husain's principles
Why Watch? The video shows the complete workflow from zero to first evaluation. Seeing it in action makes the documentation come alive and helps you understand:
- How the annotation UI works
- Keyboard shortcuts in practice
- What a failure taxonomy looks like
- How evals are configured and run
If you're developing or contributing to EmbedEval, use the embedeval-dev script:
# Clone the repository
git clone https://github.com/Algiras/embedeval.git
cd embedeval
# Check your dev environment
./embedeval-dev --doctor
# Install dependencies
./embedeval-dev --install-deps
# Build TypeScript
./embedeval-dev --build
# Run CLI commands (no global install needed)
./embedeval-dev collect examples/v2/sample-traces.jsonl
./embedeval-dev view test-traces.jsonl
./embedeval-dev annotate test-traces.jsonl --user "dev@local"
# Development utilities
./embedeval-dev --watch # Watch mode for auto-rebuild
./embedeval-dev --test # Run test suite
./embedeval-dev --lint # Run ESLint
./embedeval-dev --types # TypeScript type check
./embedeval-dev --clean # Clean build artifacts# Import traces from JSONL
embedeval collect ./logs.jsonl --output traces.jsonl
# Interactive annotation (p=pass, f=fail, s=save)
embedeval annotate traces.jsonl --user "pm@company.com"
# Read-only viewer
embedeval view traces.jsonl
# Build failure taxonomy
embedeval taxonomy build --user "pm@company.com"
# Display taxonomy
embedeval taxonomy show# Add evaluator (interactive wizard)
embedeval eval add
# List evaluators
embedeval eval list
# Run evaluations
embedeval eval run traces.jsonl --config evals.yaml
# Generate report
embedeval eval report --results results.jsonl# Create dimensions template
embedeval generate init
# Generate synthetic traces
embedeval generate create --dimensions dims.yaml --count 50
# Export to Jupyter notebook
embedeval export traces.jsonl --format notebook
# Generate HTML dashboard
embedeval report --traces traces.jsonl --annotations annotations.jsonl# GOOD: Clear, fast decisions
evals:
- name: is_accurate
type: llm-judge
binary: true # Only PASS or FAIL
# BAD: Never do this
evals:
- name: quality_score
type: 1_to_5 # Creates disagreement# Spend 60-80% of time here:
embedeval annotate traces.jsonl --user "expert@company.com"
# NOT here (automate only after understanding):
# embedeval eval run traces.jsonl # (do this AFTER annotation)evals:
# Run cheap evals first
- name: has_content
type: assertion
check: "response.length > 100"
priority: cheap
# Expensive evals only for complex cases
- name: factual_accuracy
type: llm-judge
priority: expensive# One "benevolent dictator" owns quality:
embedeval annotate traces.jsonl --user "product-manager@company.com"
# Not multiple people voting (causes conflict)npm install -g embedevalnpx embedeval collect ./logs.jsonlgit clone https://github.com/Algiras/embedeval.git
cd embedeval
npm install
npm run build
npm link # Makes 'embedeval' command available globallyOne JSON object per line:
{"id": "trace-001", "timestamp": "2026-01-30T10:00:00Z", "query": "What's your refund policy?", "response": "We offer full refunds within 30 days...", "metadata": {"provider": "google", "model": "gemini-1.5-flash", "latency": 180}}{"id": "ann-001", "traceId": "trace-001", "annotator": "pm@company.com", "timestamp": "2026-01-30T10:05:00Z", "label": "fail", "failureCategory": "hallucination", "notes": "Made up refund time limit"}evals:
- id: has_content
type: assertion
priority: cheap
config:
check: "response.length > 50"
- id: accurate
type: llm-judge
priority: expensive
config:
model: gemini-1.5-flash
prompt: "PASS or FAIL: Is this accurate?"
binary: true# Collect week's traces
embedeval collect ./logs/week-$(date +%Y-%m-%d).jsonl
# Sample 100 for annotation
head -n 100 traces.jsonl > sample.jsonl
# Annotate
embedeval annotate sample.jsonl --user "pm@company.com"
# Build/update taxonomy
embedeval taxonomy update
# Run all evals
embedeval eval run traces.jsonl --config evals.yaml
# Generate report
embedeval report --traces traces.jsonl --annotations annotations.jsonl# 1. Build taxonomy to see top failures
embedeval taxonomy build
# 2. Add eval for top category (e.g., hallucination)
embedeval eval add
# Interactive wizard asks for type, model, prompt
# 3. Run the new eval
embedeval eval run traces.jsonl --config evals.yaml# 1. Create dimensions file
embedeval generate init
# Edit dimensions.yaml to define test scenarios
# 2. Generate synthetic traces
embedeval generate create -d dimensions.yaml -n 50
# 3. Run your system on synthetic queries
# (Implementation depends on your system)
# 4. Evaluate
embedeval annotate synthetic-traces.jsonl --user "tester@company.com"For Claude, Cursor, or other MCP clients:
{
"mcpServers": {
"embedeval": {
"command": "npx",
"args": ["embedeval", "mcp-server"],
"env": {
"GEMINI_API_KEY": "your-api-key"
}
}
}
}See LLM.md for detailed agent usage guide.
Already configured. Site updates automatically on push to main.
GitHub Actions workflow included. Runs on every PR:
- Type checking
- Build verification
- CLI command tests
- LLM.md - Complete guide for AI agents and LLMs
- GitHub Pages - Visual documentation
- Hamel's Eval FAQ - Methodology reference
- β Interactive Annotation - Terminal UI for fast binary annotation
- β Failure Taxonomy - Auto-categorize failures with axial coding
- β Binary Evaluation - Assertions, regex, code, LLM-as-judge
- β Synthetic Data - Dimension-based test generation
- β Jupyter Export - Statistical analysis notebooks
- β HTML Reports - Shareable dashboards
- β JSONL Storage - Simple, grep-friendly, no databases
- β Zero Infrastructure - No Redis, no queues, no setup
v1 had 88 files with complex A/B testing, genetic algorithms, and BullMQ queues.
v2 has ~20 files with a simple philosophy: look at your traces first.
Before: Infrastructure-heavy, hard to understand
After: Simple CLI, clear workflow, Hamel Husain principles
MIT
Built with β€οΈ following Hamel Husain's principles. The goal is understanding failures, not perfect metrics. Spend time looking at traces! π