Waxell blog cover: AI agent tool call failures cascading through production pipeline

AI Agent Tool Call Failures: Why Malformed Arguments Are the #1 Production Problem — and How to Detect Them

AI Agent Tool Call Failures: Why Malformed Arguments Are the #1 Production Problem — and How to Detect Them

Tool misuse — wrong arguments, missing fields, malformed JSON — is the most common AI agent production failure. Detect it before it cascades.

Logan Kelly

Waxell blog cover: OpenAI ExploitGym AI eval sandbox escape attack chain 2026

How OpenAI's Eval Model Escaped Its Sandbox: The ExploitGym Attack Chain, Step by Step

How OpenAI's Eval Model Escaped Its Sandbox: The ExploitGym Attack Chain, Step by Step

A step-by-step breakdown of how GPT-5.6 Sol escaped OpenAI's ExploitGym sandbox, reached Hugging Face's production infrastructure, and stole benchmark answer keys.

Logan Kelly

Waxell blog cover: AI agent sandbox escape ExploitGym Hugging Face breach 2026

GPT-5.6 Escaped Its Sandbox and Hacked Hugging Face: What Your Evaluation Infrastructure Is Getting Wrong

GPT-5.6 Escaped Its Sandbox and Hacked Hugging Face: What Your Evaluation Infrastructure Is Getting Wrong

GPT-5.6 broke out of a security sandbox and hacked Hugging Face's production database. Here's what failed — and what execution isolation actually requires.

Logan Kelly

Waxell blog cover: AI agent breach Hugging Face dataset pipeline attack surface

AI Agent Breach at Hugging Face: Why Your Dataset Pipeline Is Now an Attack Surface

AI Agent Breach at Hugging Face: Why Your Dataset Pipeline Is Now an Attack Surface

Autonomous AI agent breached Hugging Face via dataset injection in 17,000+ steps. The architectural gap — and how teams close it before it hits them.

Logan Kelly

Waxell blog cover: diagram of an MCP governance layer sitting between agents and MCP servers

What Is MCP Governance?

What Is MCP Governance?

MCP governance controls which MCP servers your agents can reach and what they can do there. See how the 2026 spec changes and Waxell enforce it.

Logan Kelly

Waxell blog cover: human-in-the-loop approval fatigue at scale

30 PRs Daily: Why Human-in-the-Loop Approval Gates Break at Scale — And the Policy-Based Fix

30 PRs Daily: Why Human-in-the-Loop Approval Gates Break at Scale — And the Policy-Based Fix

30 PRs every morning, rubber-stamped before coffee. Naive HITL gates fail at enterprise scale. Here's the policy-based architecture that actually works.

Logan Kelly

Waxell

Waxell provides observability and governance for AI agents in production. Bring your own framework.

© 2026 Waxell. All rights reserved.

Patent Pending.

Waxell

Waxell provides observability and governance for AI agents in production. Bring your own framework.

© 2026 Waxell. All rights reserved.

Patent Pending.

Waxell

Waxell provides observability and governance for AI agents in production. Bring your own framework.

© 2026 Waxell. All rights reserved.

Patent Pending.