Blog
Notes from the build.
Deep dives into how the runtime actually works, covering the design decisions and the plumbing behind the product. Mostly written by the people who wrote the code.
Recurring automation, measured.
This blog keeps claiming the architecture is cheaper and steadier on recurring work, so we built a benchmark suite that hands the identical English sentence to us and to three other open-source agents, and wrote down what the ledgers said. Losses included.
Two brains, one voice.
A model smart enough to do the work is too slow to hold a conversation, and a model fast enough to hold a conversation isn't smart enough to do the work. So we run two at once. Here's roughly every decision that went into stopping the caller hearing the seam.
Why we split skills into functions and guidance.
Skill folders nest scripts inside prose. We keep two libraries instead, executable functions and prose guidance linked many-to-many, and I think it changes what an agent can actually reuse.
The conversation is not the work loop.
On the layer that sits between you and the agents doing the work, why most frameworks don't have one, and why Thinking Machines recently argued they should.
Reading about it is the slow way round.
Starter credits, no card. Hand a droid one real job and watch what it does with it.