One-person applied R&D practice agent systems, formal methods, evidence

I build agent systems that can prove what they did.

This site explains the work. The public agent-skills repository holds the source.

I take on hard-to-staff agent, compliance, and multimodal R&D and deliver working source, deterministic checks, runbooks, receipts, and explicit limits your team can inspect and own.

An unusual résumé: commercial composer for Adidas and Pepsi, Webby-recognized producer for Sony, DARPA technical lead alongside Lockheed Martin and MIT. High-end creative and hard technical work, delivered by the same person, shipped as working code, in public.

Horus Lupercal & Embry, taking tea in a dream

Not a keyframe: a rendered persona-dream. Embry dreams a quiet argument with Horus about building SPARTA Explorer as a campaign of proof: no unsupported claim past the perimeter, no receipt pretending to be a verdict. The pipeline dreams it from her own memory residue, checks it's still her, and writes it back to memory — the source and evidence state are in Explore.


Tau Selected investigations

One dominant proof, three supporting systems.

preview here → full index one step deeper

Agentic pipelines & orchestration50 skills

Reliable agent execution and orchestration — contracts, transport, and runtime truth.

surft'auscillmdebugger

Agentic memory & persona28 skills

Durable memory, persona, and voice an agent keeps and re-queries as its own history.

persona-dream

Extraction & evidence67 skills

Documents, research, and video into one truthful, cited result — no knobs to mislead.

extractordogpilewatch

Compliance, security & governance59 skills

Evidence-grounded reasoning where a human, not the model, holds authority.

sparta explorer

Adaptive-lineage hacking1 skill

Adversarial evolution and security testing scored by deterministic proof gates.

battle

Most current work is export-controlled or sensitive. The public pattern here is deliberately narrower: problem class, public artifact, what it proves, and what it does not prove. Private systems stay private.

Public
Owned/open skills, receipts, proof dossiers, and synthetic or product-owned visuals.
Private
Client names, program details, data, screenshots, private repos, and evidence counts.
Handoff
Source, checks, runbooks, and boundary notes delivered inside the client stack where practical.

01 Dominant investigation

How do you build an agent harness that doesn't trust agents?

Tau's answer: a contract around every turn — memory-first context before action, bounded subagents, goal-locked handoffs, explicit mocked/live proof boundaries. No receipt, no action.

  1. GoalContain agent work with a goal, evidence, receipt, and stop condition.
  2. Bounded dispatchSubagents run under contracts instead of open-ended delegation.
  3. Captured executionThe public sanity receipt records exit code 0 and 45 passed.
  4. BoundaryThe receipt names browser UI and live-provider semantics as unproved.
  5. Receiptskills/tau/proofs/p0-operator-wrapper-20260705T1200Z/sanity.json

What remains unproved: equal proof depth across every skill and fresh live-provider behavior.

Tau sanity receipt excerpt showing exit code 0, 45 passed, and explicit proof boundaries.t'auMemory-first zero-trust harness: every agent turn gets a goal, evidence, receipt, and a stop condition.
01

sparta explorerspace-cyber evidence workbench

Who decides, when the AI sounds sure? Humans do.

Space-cyber evidence threads traced to human decisions — relevance never becomes authority.

Evidence threads traced from framework guidance to program evidence to a human decision — relevance never becomes authority, candidate coverage never becomes credit.

external repo
02

persona-dreamdream-affect voice study

Can a dream give a persona more than memory can?

Preregistered study: dreams as inspectable intermediates whose certified affect drives the persona’s live voice — emotion tags, tone, identity held stable.

A preregistered study: the dream is an inspectable intermediate whose certified affect drives the live Chatterbox voice — emotion tags, tone — and “No” is a real answer.

sanity-checked
03

battleexploit-evolution arena

What do you do with exploit code you can't trust? Breed it.

Red/blue exploit evolution in isolated Docker targets, scored by deterministic proof gates.

Red and Blue evolve attacks and defenses in isolated Docker targets; generated code is genetic material, and a deterministic Judge — not plausibility — scores what actually worked.

sanity-checked

Ledger agent-skills source map

352 public contracts; 295 with sanity checks.

The full inventory is still public, including contract-only entries and visible gaps. It now lives where a technical inspector expects it: on a direct, reloadable route.

Inspect all contracts

Tau How proof works

One real run, from goal to receipt.

Not a diagram of an idealised pipeline: the actual tau roundtable that designed this page, walked stage by stage. Each step resolves to a real immutable artifact you can hash-check, and each one says plainly what it does not prove.

Open the proof route

Receipts Bounded evidence

No claim ships without one.

2 captured excerpts, printed as they came out of gen_artifacts.py: a node receipt from the roundtable run that designed this page, a captured audit, and the provenance of the numbers above. Missing sources stay visible as boundaries, not status widgets.

tau node receipt — the run that designed this site

PREFLIGHT: PASS

A real model produced this seat's response, on the record, at a stamped time.

That the response's content is right — only that it genuinely occurred.

node-receipt.json · ask-tau-roundtable-round-2-webgpt-webcla-ce79c8ef960c · sha256 e697899eb7dcc021…. The judgment, proof boundary, and raw payload stay together; the summary is never the only evidence.

node-receipt.json · ask-tau-roundtable-round-2-webgpt-webcla-ce79c8ef960c · sha256 e697899eb7dcc021…

monitor-website audit — this page, checked against the repo

UNAVAILABLE

{ "drift": [ "stats.skills: README=347 site=352", "stats.sanity: README=292 site=295" ], "live": { "home": { "status": 200, "ok": true }, "sitemap": { "status": 200, "ok": true }, "resume": { "status": 200, "ok": true }, "resume_pdf": { "status": 200, "ok": true }, "resume_md": { "status": 200, "ok": true } } }

That local generated surfaces agreed or public endpoints responded.

inventory provenance — where the numbers come from

BUILD: 1049d151

The numbers above came from checked source state, not marketing copy.

That every skill is complete, useful, or production-ready.

site/inventory.json · sha256 84864adaf4ababc4…. The judgment, proof boundary, and raw payload stay together; the summary is never the only evidence.

site/inventory.json · sha256 84864adaf4ababc4…


Person Accountability

An unusual path, on purpose.Composer — Adidas, Pepsi, X-GamesExecutive producer, SonyDARPA ARCOS — principal data scientistAFRL “Hacker” challenge coinLean 4 formal methodsThis practice

An unconventional path is an advantage on problems with no playbook.

One person also means direct accountability — the person you meet is the person who investigates, architects, builds, and answers for the result. Available for engagements and full-time roles — full résumé.

Composer

Commercial work for Adidas, Pepsi, X-Games.

Executive producer, Sony

God of War: Ascension campaign — Webby-recognized, 80-person productions.

DARPA ARCOS

Principal data scientist and technical lead, alongside Honeywell, Lockheed Martin, MIT, GE, SRI.

AFRL “Hacker” challenge coin

Recognition out of that work.

Lean 4 formal methods

Proof discipline carried into agent design.

This practice

Agent systems that produce their own evidence — shipped as working code, in public.

Next Evidence-first work

Bring me the project you shelved.

The one with no playbook — the one that stalled because it needed both halves of the job. Good fits:

  • Stalled multi-agent orchestration — pipelines that demo well but fail silently or drift in production.
  • Zero-trust & compliance blockers — work that can't ship without audit trails, receipts, and evidence.
  • Platform-independent R&D — you need the work done inside your stack, delivered as open code you own, with no agency overhead and no vendor platform to adopt.
  • Bespoke multimodal & generative workflows — where the work demands technical rigor and aesthetic polish at once.

Principal-led, not principal-dependent: the deliverable is source your team owns, deterministic checks, runbooks, and plain boundary notes so another engineer can continue the work.

graham@grahama.co