The Series
Field notes from a daily practice with AI coding agents. Patterns, tools, sharp edges. Evidence-based, no fluff.
The Score I Never Measured
I wrote that a prompt re-scored against three test cases, all passing. I had run zero of them, inside a document about unverified claims.
The Step That Never Ran
The job died installing dependencies, so the CLI check never executed. Two verifiers reported totals 5.1 MB apart, and they were reading two different job logs.
Two Hundred OK With Five Missing
A negative limit returned a success status and ok:true while the advertised count and the payload disagreed. The envelope reported success over a silent drop.
The Guard That Announced Its Own Absence
A bash guard walked every file in the repo on each call, overflowed its own match cap, and printed that its coverage was OFF dozens of times in one session.
What we ship
ValidationForge proves that code works. Anneal proves the plan works before any code gets written. Two flagships, both production-grade. The other cards are companion tooling from the series.
Get the next field note before it ships.
4,206 engineers, operators, and weirdos. No spam, ever.