- San Francisco
- https://unhype.me
- @esthor
- in/erikthorelli
Highlights
evals.ax
🧊 Open source LLM observability platform. One line of code to monitor, evaluate, and experiment. YC W23 🍓
🤘 TT-NN operator library, and TT-Metalium low level kernel programming model.
Holistic Evaluation of Language Models (HELM) is an open source Python framework created by the Center for Research on Foundation Models (CRFM) at Stanford for holistic, reproducible and transparen…
A framework for few-shot evaluation of language models.
A framework for few-shot evaluation of autoregressive language models.
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
The open source research environment for AI researchers to seamlessly train, evaluate, and scale models from local hardware to GPU clusters.
An MCP server that autonomously evaluates web applications.
UI Library for Design Engineers. Animated components and effects you can copy and paste into your apps. Free. Open Source.
Model Context Protocol Servers
Fully local web research and report writing assistant
Enterprise grade context intelligence for everyone.
Benchmark environment for evaluating vision-language models (VLMs) on popular video games!
Stop configuring your AI stack. Start using it. One command brings a complete pre-wired LLM stack with hundreds of services to explore.
AdalFlow: The library to build & auto-optimize LLM applications.
OpenTelemetry Instrumentation for AI Observability
Our library for RL environments + evals
Agent2Agent (A2A) is an open protocol enabling communication and interoperability between opaque agentic applications.
How Python does AI. Agents, realtime voice, image generation, embeddings. Every model, every interface, typed end to end.
The SOTA Open-Source Browser Agent for autonomously performing complex tasks on the web
Laminar - open-source observability platform purpose-built for AI agents. YC S24.
Langtrace 🔍 is an open-source, Open Telemetry based end-to-end observability tool for LLM applications, providing real-time tracing, evaluations and metrics for popular LLMs, LLM frameworks, vector…
The AI Toolkit for TypeScript. From the creators of Next.js, the AI SDK is a free open-source library for building AI-powered applications and agents
Inspect: A framework for large language model evaluations