Menu
Production AI · eval-gated · yours at handoff

Most AI demos die in production. We build the services that don't.

We're the team you bring in when the prototype impressed everyone and now has to survive real users, real queues, and real operating cost. RAG, agents, and voice — shipped with the evals, observability, and runbooks that tell you whether the system should scale, and a go/no-go recommendation in writing before you commit.

View work
Eval ledger LIVE
#1042 faithfulness 0.94 281ms PASS
#1041 injection-suite 0.98 PASS
#1040 groundedness 0.87 302ms FAIL
#1039 refusal-bounds 0.96 PASS
#1038 cost-ceiling 0.99 274ms PASS
#1037 faithfulness 0.93 288ms PASS
CI-GATED · MAIN DEPLOY UNLOCKED
  • ENGAGEMENT
    06– 14 WK

    Typical engagement window, from diagnostic to handoff.

  • SERVICES
    06

    Six ways to retire production risk: RAG, agents, voice, post-training, agent security, and payment rails.

  • TEAM SIZE
    02 ENG

    Senior engineers embedded inside your repo, with one accountable path to deploy.

  • SUPPORT
    30 DAY

    Post-handoff support window while ownership moves fully to your team.

PRODUCTION STACK
Claude · GPT-5 · Anthropic MCP · LangGraph · pgvector · bge-m3 · Cohere Rerank · Modal · Temporal · LiveKit · Deepgram · ElevenLabs · x402 · ERC-8004 · AP2 · Foundry · ElizaOS · Claude · GPT-5 · Anthropic MCP · LangGraph · pgvector · bge-m3 · Cohere Rerank · Modal · Temporal · LiveKit · Deepgram · ElevenLabs · x402 · ERC-8004 · AP2 · Foundry · ElizaOS ·
01 / SERVICES

Six services. Every one built to hold up in production.

01 / SERVICE

RAG that holds up under eval

Retrieval-augmented generation with the eval harness built in — an index is the right call on layout-heavy, citation-bound, or latency-tight work; the diagnostic makes that call before the build does, not an assumption baked in on day one. We pick the chunker, the embedding model, and the retriever — contextual chunking and late-interaction retrieval where they earn their place — and benchmark every change against a golden set you'll keep using after we leave. The goal is fewer unsupported answers, less manual review, and degradation caught before customers see it. The same harness ships on its own, for a model already in production with no way to know when it degrades.

  • Hybrid + late-interaction retrieval, agentic re-query
  • Contextual & late chunking per document class
  • ColPali visual retrieval for layout-heavy docs
  • Faithfulness + groundedness evals, CI-gated
pgvector · bge-m3 · cohere-rerank
02 / APPROACH

An engagement looks like this — predictable by design.

W1–2 STEP 01

Diagnostic

We sit with the team, read the data, and retire the riskiest assumption first. The output is a 12-page memo: what to build, what to skip, what eval to point it at.

2 weeks fixed
1 principal eng
Memo + spike repo
W3–10 STEP 02

Build

We pair with your engineers in-repo. Eval gates from day one, weekly demos against the metrics defined in week 2, and tradeoffs made while the code is still cheap to change.

6–8 weeks
2 senior engs
Production deploy
W11+ STEP 03

Handoff

Runbooks, eval suite, dashboards, on-call rotation, and a 30-day support window. The deliverable is not just code; it is your team's ability to run it without us.

2 weeks fixed
Docs + runbooks
30 day support
03 / METHODOLOGY

What you'll actually have, week by week.

W1–2

Diagnostic

INPUTS WE NEED
  • A representative data sample.
  • Your current eval suite — even if it's a sheet.
  • One engineer, full attention, for week one.
WE SHIP
  • A 12-page memo: what to build, what to skip.
  • A spike repo with the riskiest path proven out.
  • An eval-suite skeleton, wired and runnable.
WE MEASURE
  • Open questions closed by end of week 2.
  • The riskiest assumption retired or named.
  • Decision clarity — yes / no / not yet.
W3–10

Build

INPUTS WE NEED
  • Repo access and a CI lane we can break.
  • Authority to decide tradeoffs in real time.
  • A 45-minute demo slot, weekly, no slides.
WE SHIP
  • Production deploy gated by the eval suite.
  • A dashboard for eval-pass rate and p95.
  • Runbook v1 — incidents, rollback, scaling.
WE MEASURE
  • p95 latency against the budget set in week 2.
  • Eval-pass rate, run-over-run.
  • Deploy confidence — gated, observable, reversible.
W11+

Handoff

INPUTS WE NEED
  • Your on-call rotation and pager policy.
  • The team that will own this, named, not TBD.
  • Two half-day training sessions on the calendar.
WE SHIP
  • Runbooks, eval suite, dashboards — yours.
  • On-call rotation handover with shadow shifts.
  • 30 days of on-tap support, no scope haggling.
WE MEASURE
  • Incidents resolved without us.
  • MTTR — pre vs. post handoff.
  • Eval-suite coverage your team can extend.
04 / WHAT WE BUILD

Eight representative engagements.

Representative engagements — the problem each one starts from, and the system we build to solve it. Examples span financial services, manufacturing, and construction. The examples are qualitative, not a claim of delivered client results.

01 / 08
01 ENGAGEMENT · Insurance

Voice intake for first-notice-of-loss claims

VoiceRAGAgents
PROBLEM

Long IVR handle times and low first-call resolution. Every triage minute is a customer thinking about switching.

RISK

Claims leakage starts as queue pressure: slow intake, inconsistent routing, and avoidable escalations that look like service quality problems.

SYSTEM

Voice agent architected to a tight real-time latency budget, with hybrid retrieval against policy documents and human-in-the-loop review on liability calls.

Architecture Policy docs · liability fork
01
CALLER Inbound FNOL call
02 p95 280ms
VOICE AGENT STT → LLM → TTS, barge-in
03 0.94 faithful
HYBRID RAG Retrieval over policy docs
04
RESOLVE Auto-settled claim
05
HITL Liability review
06 / FAQ

Questions we answer on the first call anyway.

How do you charge?
Fixed fee for the diagnostic — two weeks, paid up front. The build phase is a weekly retainer, scoped to the deliverables set in week two. We bill outcomes, not hours. No success fees, no equity, no kickers.
Do you sign NDAs?
Yes. Mutual NDA — our paper or yours — signed before the first technical conversation. Most teams send their own; we counter-sign within a business day.
What does “eval-gated” actually mean in practice?
Every commit runs against a versioned eval suite. If the eval-pass rate drops below the budget set in week two, the deploy doesn't ship. The suite is yours at handoff — runner, dataset, scoring rubric, and the dashboard that watches it.
Will you push us off our existing stack?
No. We work with the models, vector stores, and frameworks you've already chosen — unless one of them is the reason the project is stuck. If so, the diagnostic memo says so, and you decide.
Who owns the IP and the eval suite after handoff?
You do. All of it. Code, evals, runbooks, dashboards. We retain no rights, no licenses, no required attribution. The only thing we keep is the right to reference the engagement publicly — with your written sign-off.
What if the diagnostic recommends not building?
You keep the diagnostic memo and the analysis behind it, and you make the call with clear eyes. A sound go/no-go decision is a real outcome of the engagement — getting that call right matters as much as shipping the system itself.
How do we know this is worth building?
The diagnostic starts there. We identify the workflow, the failure mode, the user who owns it, and the eval that would prove the system is improving the work rather than adding another tool to babysit.
What does our team need to own after handoff?
The repo, the eval suite, the dashboards, the runbooks, and the on-call path. We do not hand over a black box; we hand over the operating surface your engineers need to debug, extend, and retire parts of the system when the workflow changes.
Can you work with business stakeholders as well as engineering?
Yes. Engineering owns the system, but the workflow usually belongs to ops, risk, support, sales, or finance. We keep the technical interface precise and translate the build/no-build decision into the operational risk it retires.
How fast can you start?
Diagnostic phase usually starts two to four weeks after the first call. We run one diagnostic at a time, so the calendar is the constraint. Right now we're booking into Q4 2026.
◇ NEXTA 30-MIN CALL · NO DECKS

Bring us a hard problem. We'll show you what we'd build.

The first call is a free 30 minutes. You'll leave with the first cut of the build/no-build path, the riskiest assumption, and the eval we'd use to test it.

or hello@proofoftech.org
NEW ENGAGEMENT · INTAKE

Tell us about it.

The more specific you are, the more useful our first reply.

SERVICE AREA
↩ ENCRYPTED IN TRANSIT
ASK THE FIELD NOTES BETA