A running notebook of experiments, half-built ideas, and things I learned from poking at session data. New notes land here first. Subscribe via RSS
-
Skills tastemaker
We make skills easily and reuse them rarely, so I took my last 100 Claude Code sessions and asked which ones actually deserve to become one. About 30 had a clear, repeatable procedure. Setup and release work turned into skills again and again. Debugging never did. The funniest tell so far is that a proper noun in a skill's name usually means it's junk.
-
34 seconds of paperctl labels
I made a terminal demo of the paperctl label workflow: create a label, tag sessions, filter the fleet by it, export an audit trail. Five commands, 34 seconds, rendered frame by frame so the typography could do things a live terminal can't. Every command echo is verified against the actual CLI source, and every session in it is fabricated. The hard part of a good terminal demo is that real terminals are ugly at 2560 by 1440.
-
Does the expensive model save you work?
Two months of team session data, roughly 600 real working sessions, one question: does paying more per turn reduce the human work of supervising an agent? Paying 2x for a same-generation model changed nothing we could measure; corrections, interrupts, and rework all landed in the same range. Upgrading one model generation at the same price dropped fix-it follow-ups from 21% of sessions to 2%. So the generation axis moves your workload and the price axis mostly moves your bill. Bonus lesson: regexing for 'no, actually' is a bad correction detector, because behavior signals like rework beat anything that depends on phrasing.
-
My quarterly review, from session data
I did my quarterly review from my own session capture instead of memory, and the data called me a liar immediately: 351 sessions became 126 once I separated real working sessions from noise. A fifth of my model calls were permission checks (I had zero allow rules configured while a skill for exactly that sat installed), I re-read the same context files 19 times, and I rebuilt the same export-and-classify pipeline 9 separate times without ever writing the skill. I've written about 35 skills this year; my most-reused artifact is a plain markdown file. Adoption follows friction, and apparently I'm not exempt.
-
Put the model on trial
We gave a frontier model two weeks of real engineering work, then staged an actual trial: prosecution and defense subagents arguing from the session records, a judge striking every claim the evidence couldn't support. The judge struck five, and three of them were us blaming the model for our own environment. The real verdict was that most of the model's work couldn't be judged at all, because an export failure ate the records. You can't grade a model on evidence you didn't keep, so fix the telemetry before you fix the model.
-
Onboarding, according to the session logs
I mined 92 days of team session records, about 2,900 sessions, to see what ramping up actually looks like, measuring time-to-first-PR straight from the gh pr create calls in the trace. Ramp has a visible shape: first an environment gauntlet, then a long stretch of learning the codebase's conventions by failing against them. The embarrassing part is that my 'new hire' turned out to have commits from months before my data window started, so I was really watching an experienced teammate ramp on a new workflow. Session data will happily let you tell yourself a story if you don't check the edges of the window.
-
Session ghosts
Three engineers investigated the same slowdown four times over 29 days and never saw each other's work: seven sessions, 632 tool calls, and the same handler file independently opened three times, 25 days apart. So I built a Mario Kart ghost overlay: past sessions rendered translucent alongside the current one, with the ten files where trails converge glowing. The data was already sitting in the fleet export; the visualization is a 90-line zone classifier and an HTML render. Your teammates' dead ends are the cheapest map you'll ever get.
-
The architecture decision stands trial
We took a month-long migration we'd already shipped and put the decision itself on trial: a prosecution subagent, a defense subagent, and a judge who checks every claim against 70 recorded sessions and strikes anything the record can't support. The judge struck six claims, from both sides. The verdict came back split: right call on the outcome, unprovable on the process, because the record contains no weighing of alternatives. Nobody gets to invent evidence when every citation has to quote a session character for character.
-
Paper discovers itself
We pointed paper at our own fleet: three weeks of session records, run through an adversarial refute-everything review. The budget mostly goes to re-reading. Input tokens outnumbered output 375 to 1, one in four model calls was a permission check re-reading a huge cached prefix, and the single most common prompt in the whole fleet was the harness asking 'where were we?'. Also, there is no average session; a handful of whales carry most of the spend. Turns out the agent's biggest expense is remembering what it was doing.
-
Reading the cache like a diagnostic
Three SQL queries against my own archive: 93% of input tokens came from cache, which puts effective input cost at roughly a sixth of nominal. The per-model split was the interesting part. Opus hit 95.6%, Sonnet 94.4%, Haiku 58.6%, and that gap says more about how each model gets used than about the models themselves. Cache hit rate turned out to be a prompt-stability metric hiding in the billing data.
-
The average workday doesn't exist either
Same archive, different question: when does agent work actually happen? Five of thirteen active days carried 92% of all messages, and two days alone carried 58.5%. Sprint days have a shape too: peaks at 10am, noon, and 2pm, a hard crash at 4pm (two messages), a partial recovery at 8pm. Monthly dashboards smooth all of this away. Token leaderboards are shaping up to be the new lines-of-code metric, and they'll reward exactly the wrong engineer.