StellarRequiem · Security Research & Verification Infrastructure

Alex Price / StellarRequiem

Verified work, or it doesn't ship.

I build security and verification tooling for AI-era work: MCP authorization research, scanner benchmarks, public-safe workflows, and proof-carrying claims. The dangerous outputs are the confident-sounding numbers nobody verified. One rule runs through everything I ship: no belief without verification — every result carries evidence a third party can re-run. The badge is the claim; the honest no is part of the deliverable.

FastMCP fixes 2 merged MCP benchmark 0/11 MCP proofs public Workflow public Proof re-runnable Honest gaps stated

Responsible Security Research

Offensive technique applied under explicit authorization, with audit. Recent work centers on MCP / AI-infrastructure boundary failures: versioned authorization checks, replay/session isolation, and scanner blind spots around authorization logic.

Authorization framework · public

scope-gate

A deny-by-default authorization gate: test only what you're explicitly authorized to. Ships with a responsible-research charter. The boundary that makes dual-use work safe.

PublicDeny-by-default
Public upstream fixes

FastMCP merged fixes

Two security-relevant MCP boundary issues I identified were fixed upstream: versioned authorization checks and Streamable HTTP event replay isolation. Cited as public merged fixes, not CVE/GHSA claims.

2 merged PRsRegression coverageNo advisory claim
Re-runnable fixtures · public

MCP authorization proofs

Local-only reference fixtures for MCP-shaped authz boundaries: resource/audience/scope binding, session replay, token handling, and a FastMCP signed-agent path with explicit allow + deny cases. Numbers and limits live on the proof page — fixture mechanics only, not production audit or conformance.

Re-runnableDeny paths explicitClaim boundaries on page
Discipline

Audit everything

Append-only run journals, claim cards, and hash-chain receipts keep the work grounded: what changed, what is verified, what is only a lead, and what still needs an operator gate.

Append-onlyClaim cardsHash-chained
Benchmark · public

mcp-bench

Do MCP security scanners actually catch authorization-logic bugs? An independent, reproducible benchmark seeded with real confirmed findings: 19 labeled cases, including 11 authz-logic cases across 10 root-cause classes, 2 control bugs, and 6 clean negatives. Scanners run only in a disposable CI runner.

PublicReproducibleauthz-logic 0/11
vulnerability researchPoC development authorization-gated testingcoordinated disclosure MCP / AI-infra securityFastMCP public fixes claim cardsred-team tooling

Verified AI Labor — the platform

Can a company be run as agents? Only if you can trust what each agent says it did — so I built the loop around verification, scope gates, false-positive rails, and public-safe workflow artifacts.

The gate

verity-core

Refuses a "95% accuracy" claim until it clears statistical hygiene — sample, out-of-sample, leakage, lift over base rate — then proves it: a claim ships a re-runnable command and the number must reproduce or CI fails. 17 domain packs · CI gate · MCP tool.

Proof-carryingCI gate
The labor

verified-ai-labor

A working prototype of a company run as agents — a 13-stage pipeline where every result-claim is verity-gated and every action hash-chain-logged, observable in a live console. Tests run locally; CI badge pending.

Verity-gatedHash-chained
Governed autonomy

The operating surface

Agents run a real workstation — but every action routes through a deny-by-default reference monitor first: read and local work proceeds, anything outward or destructive is held for a human, money and credentials are refused. Model proposes, code disposes — every decision hash-chained, every window operator-summoned. No capability the gate didn't grant.

Deny-by-defaultHash-chainedOperator-gated
The benchmark

groundtruth-bench

Citation faithfulness you can re-run to the same hash: a cryptographically committed corpus scored offline, byte-identical across machines — where RAG eval (RAGAS/ARES) is online, metered, and uncommittable. Reports where the scorer fails, not just the flattering number.

Byte-reproducibleCommitted corpus
The proof

calibration-log

A public, hash-chained prediction record scored over time (Brier + calibration). Honesty you can't doctor — it reports the real number whether there's an edge or not.

LiveHash-chained
The adjudicator

scorecheck

Adjudicates a published benchmark claim against its raw run-logs — REPRODUCED / DID-NOT-REPRODUCE / CHERRY-PICKED — sealed into a re-runnable receipt. Surfaces the dropped, flipped, and fabricated rows that re-run leaderboards and reproducibility badges miss; survived a 3-lens adversarial pass.

ReproducedCherry-pickedCI-green
Trust tooling

firewall · grounded · reality-anchor

Flag unverified claims in AI output; verify every cited claim is supported by its source; a research agent that grounds its answers or abstains rather than fabricate.

DeterministicOffline
Public operating manual

workflow bible

A public-safe breakdown of the full loop: scope, orient, map, draft, verify, ship, journal, and review. Copy the system without copying private targets, secrets, exploit steps, or live infrastructure details.

PublicCopyableSafety rails
Personal daemon · public demo

The Local Daemon

A client-side interactive console for catching signal, switching modes, and turning intensity into inspectable artifacts. Browser-only demo — not a remote backend and not wired to private infrastructure.

InteractiveClient-sideNo private backend

How my work is verifiable

Not a portfolio of assertions — a portfolio you can re-run.

01
Scope before action. Security work starts with explicit authorization, public-safe boundaries, and a clear operator gate.
02
Claim cards before headlines. Public claims get evidence, caveats, source links, and a disclosure-safety check.
03
Runnable proofs. Results come with the exact command a third party executes to reproduce them.
04
A public calibration log. Predictions are hash-chained and scored over time — the honesty is auditable, not asserted.
05
Honest gaps, stated. Every deliverable names what it did not verify. An unverifiable claim does not count.