-
eval-audit Public
A deterministic, zero-dependency Python tool that checks whether apparent evaluation anomalies survive simple alternative explanations
MIT License UpdatedSep 21, 2026 -
inspect_ai Public
Forked from UKGovernmentBEIS/inspect_aiInspect: A framework for large language model evaluations
Python MIT License UpdatedAug 27, 2026 -
TransformerLens Public
Forked from TransformerLensOrg/TransformerLensA library for mechanistic interpretability of GPT-style language models
Python MIT License UpdatedAug 26, 2026 -
latteries Public
Forked from thejaminator/latteriesJames' cookbook of evaluations and finetuning experiments
Python UpdatedAug 24, 2026 -
-
cot-monitors-wrong-sentences Public
When chain of thought monitors trustworthy?
MIT License UpdatedJul 24, 2026 -
-
-
-
AutoSteer Public
The Feature Hunter Agent for LLM Interpretability
-
rusty-llm-jury Public
Rust based CLI tool for estimating success rates when using LLM judges for evaluation.
-
ml-engineering-skills Public
Forked from stas00/ml-engineeringMachine Learning Engineering Open Book
-
sae-manifold Public
Forked from goodfire-ai/sae-manifoldcode for 'Do Sparse Autoencoders Capture Concept Manifolds?'
Python MIT License UpdatedMay 21, 2026 -
Jane street dormant puzzle trail runs
Python MIT License UpdatedMar 20, 2026 -
unified_agentic_llm_evals Public
Forked from evaleval/every_eval_everPython MIT License UpdatedMar 18, 2026 -
rust-agentic-skills Public
Rust Agentic Skills : It transforms general-purpose LLMs into a multi-agentic Rust engineering team via a token-efficient Brain-Tool-Context setup that dynamically loads tools and documentation, ad…
-
mojo-gpu-puzzles Public
Forked from modular/mojo-gpu-puzzlesLearn GPU Programming in Mojo🔥 by Solving Puzzles
-
opendistill-skills Public
Open-source, Agentic MLOps Skill for local, privacy-first SLM fine-tuning, using Unsloth for efficient QLoRA training and Ollama/lmstudio for deployment.
-
openresponses-python Public
Unofficial Python SDK for OpenResponses Spec examples.
-
original_performance_takehome Public
Forked from anthropics/original_performance_takehomeAnthropic's original performance take-home, now open for you to try!
Python UpdatedJan 22, 2026 -
Universal Shopping AI Assistant using Universal Commerce Protocol (UCP)
-
-
-
MedAgentBenchmark-Green Public
Evaluation Agent that evaluates an agent that runs tasks in virtual EHR workflows.
-
Medical Agent Benchmark - FHIR MCP Server
Python UpdatedJan 9, 2026 -
tau2-bench Public
Forked from sierra-research/tau2-benchτ²-Bench: Evaluating Conversational Agents in a Dual-Control Environment
Python MIT License UpdatedJan 5, 2026 -
garuda-engine Public
Institutional-grade async news arbitrage engine for NSE/BSE with WAF bypass and high-frequency execution
Python MIT License UpdatedDec 24, 2025 -
Paper2AgentWithSkills Public
Paper2Agent with Skills is featuring a Robust Adaptive Skill-Accretion (R-ASA) engine. It transforms research papers into active agents that dynamically synthesize, verify, and memorize executable …
-
agentic-ai-agent-beats Public
Forked from RDI-Foundation/agentbeats-tutorialAgentX-AgentBeats Competition
Python UpdatedDec 13, 2025 -
hf-skills-antigravity Public
Forked from huggingface/skillsPython Apache License 2.0 UpdatedDec 8, 2025