Introducing Vibe-Coding Arena
Most existing agent benchmarks evaluate only the outcome of a single execution—whether the final result passes or fails, but this fundamentally differs from how people actually use coding agents. Vibe coding is iterative and human-in-the-loop: users inspect intermediate results, refine requirements, correct mistakes, and steer agents across multiple rounds. We introduce Vibe-Coding Arena, an evaluation system that captures this process by having humans build alongside agents and conduct fine-grained evaluations of their behavior throughout the interaction.