This repo publishes one stage of a prose compiler — the stage that takes a PRD-shaped feature spec and produces a gradeable software artifact, with well-defined I/O contracts so other stages chain in front (PRD discovery) or after (deployment, observability).
The pipeline this stage implements:
PRD → design-doc (schema-tight classification + acceptance criteria + residue)
→ build-tools | compose (proxy gate authored from PRD; routed by FEATURE-SHAPE)
→ Phase 3.5 cross-family adversary on the gate (typed-acceptance: ENTAILMENT,
DISCRIMINATOR, SPECULATION, WRONG → SPECULATION carried in RESIDUE.md)
→ implement-spec (impl in workspace)
→ Phase 5 adversary on impl + RESIDUE re-type (SPECULATION → ENTAILMENT when impl
gives the speculation concrete shape)
→ bounded one-shot revision pass on ENTAILMENT (regression-guarded: revert if base
regressed or REWARD dropped)
→ grade (per-task verifier, unmodified)
Three deliverables (in durability order):
- The compiler stage —
skills/,harness/,frozen/,STANDARD_PROMPTS.md, the typed- acceptance protocol, the RESIDUE.md carry-forward, thedsrCLI. Reusable on any PRD-shaped feature spec with any coding-capable LLM that follows schema-tight prompts. The artifact that travels. - A bench measurement that validates the stage: Composer 2.5 in this scaffold on 109 of DeepSWE-113 (4 audit-v1 defectives excluded), pass@1 reported as Wilson 95% interval. The first published Composer 2.5 datapoint on the DeepSWE substrate — DeepSWE's leaderboard does NOT include Composer 2.5, likely because of the API gating named in deliverable #3.
- A methodology essay on what we discovered while trying to construct a defensible baseline. Cursor's structural prevention of independent Composer evaluation is the same model-vs-system collapse the DeepSWE audit identified, now in a different vendor stack. Pattern recognition over n=2 is suggestive, not proof — but earns the writeup as a methodology observation.
Models used. Composer 2.5 for craft + author + recon (via cursor-agent -p -f -w $WORK);
Gemini 3.5 Flash for Phase 3.5 adversary soundness lens (via direct API per gemini_api.py);
Composer also serves as Phase 3.5 breadth lens (dual-adversary). Both standard-tier; composer- 2.5-fast 6× markup is forbidden in the scored run. Per-token rates verified against published
vendor pricing 2026-05-29 — see docs/composer-2.5-review.md §F5.
DeepSWE's 113 tasks are used as a contamination-free
2026 substrate, graded by the unmodified Pier verifier.
Their leaderboard/recognition/PR are out of scope. See PREREGISTRATION.md.
| path | what |
|---|---|
PREREGISTRATION.md |
frozen-on-run methodology, at parity with our SWE-bench Pro prereg |
WORKLOG.md |
dated development + run trail |
skills/ |
design-doc, build-tools, implement-spec, verify-spec, compose — the skill files dispatched by dsr and the agents |
harness/ |
bootstrap.sh (per-token key validation), provision_oracle_ec2.sh (the May 27 oracle audit), smoke_box.sh (the EC2 box-side smoke), audit_oracle.sh + box_audit.sh (oracle support) |
harness/feature/ |
dsr.py (the feature-task driver CLI), build-tools-lessons.md, HYPOTHESIS_GRAPH.md, per-task run/<task-id>/ receipts |
external/ |
mirrored leaderboard JSON, codex-audit results, deep-swe pin |
results/ |
per-task verdict ledgers + run logs; results/smoke/box/ for the EC2 smoke receipts |
docs/ |
PROCEDURES.md, composer-2.5-review.md, COMPOSE-EVOLUTION.md, PR_DRAFT.md (staged for post-run) |
The most basic check before any model arm: does each task's own reference solution pass its own
verifier? Run on all 113 tasks (oracle agent, $0 model, one spot box, under a dollar), deep-swe
pinned at 2f0f4125. Result: 109 pass, 4 fail — langchain-request-coalescing,
narwhals-rolling-window-suite, prometheus-transactional-reload-status, skrub-duration-encoding,
each confirmed failing in isolation, cause unresolved. Full per-task verdicts in
results/oracle_audit_ec2.jsonl. The 4 defectives are excluded from the eligible denominator.
git clone https://github.com/datacurve-ai/deep-swe && cd deep-swe && git checkout 2f0f4125
uv tool install --python 3.12 datacurve-pier==0.2.0 # needs docker + Compose v2 plugin
for t in tasks/*/; do pier run -p "$t" --agent oracle --env docker; doneTwo harness arms on GPT-5.5 via codex CLI subscription on the eligible 109 of DeepSWE-113:
- Arm A — compiler stage scaffold (this repo): design-doc → build-tools/compose → Phase 3.5
dual cross-family adversary → impl → Phase 5 + revision-guard. All scaffold model calls via
codex exec -c model="gpt-5.5". - Arm B — codex CLI alone:
codex execon the task workspace, no scaffold. OpenAI's general-purpose agent harness.
Both arms are harnesses. Same model, same access method (codex CLI subscription), same substrate, same grader. Only the harness shape varies. McNemar paired delta + Wilson 95% marginals at α=0.05.
DeepSWE's published gpt-5-5 in mini-swe-agent xhigh, 4 trials, 0.7005 pass@1 is cited
qualitatively against both arms as an external reference point; not part of the paired stat.
Pacing. Subscription rate-limit tolerant, no API fallback. ~3000-4000 codex calls total spread over 1-3 days. Total scored-run cost: ~$25-50 EC2 + ~$1 Composer adversary calls.
Earlier framing dropped (timeline):
- 2026-05-27 original: harness-richness vs Sonnet+GPT-5.5 baselines, three arms.
- 2026-05-28: switch to Composer+Flash pair (~10× cheaper).
- 2026-05-30 first: drop Flash baseline (gemini-family generator gap structural).
- 2026-05-30 second: switch Composer baseline to mini-swe-agent (leaderboard-comparable).
- 2026-05-30 third: discovered Cursor's API gating prevents Composer-in-mini; drop the comparison entirely.
- 2026-05-30 fourth (current): switch BOTH arms to GPT-5.5 via codex CLI subscription; harness-vs-harness paired comparison, model held constant. The methodology essay grows by one observation: "every benchmark number is a harness number; model labels are a useful fiction" — see §0 deliverable #3.
| component | how | receipt |
|---|---|---|
| Compose stage end-to-end on EC2 | scaffold REWARD 1 on kysely-window-grouping-helpers via fresh spot m7i.xlarge | results/coordinator/test-drive-v2/runs/kysely-window-grouping-helpers/scaffold/ |
| Dual-adversary at Phase 3.5 | n=2 substrate measurement (bandit 37.9% overlap, kysely 11.5%) | harness/feature/run/<task>/transfer-risk-v1/ |
| Composer-as-recon dominates Flash-as-recon (n=3 head-to-head) | schema-tight prompt validation | results/recon-comparison/ |
| Multi-box coordinator dispatch | 2 concurrent spot boxes, ledger-based resume, scp pull, teardown clean | results/smoke/multibox/, results/coordinator/test-drive-v[12]/ |
| Six bugs caught + fixed by test-drive validation | cursor-agent --workspace, python3-dsr symlink, git pathspec :!, scp -r nesting, multibox keypair collision, classifier '0\n0' grep | commits 52c070b, 6f226fd, 1ea7148, 2aa0b96 |
Frozen tag: frozen-skills-v2 at commit 1ea7148. The skill collection is what's frozen; the
scored run (when launched) reads from this tag's hash manifest and refuses to start on drift.
Composer 2.5 living review with confidence-ranked findings + receipts:
docs/composer-2.5-review.md.
# defaults: DEEP_SWE_DIR=../deep-swe DEEPSWE_RUN_DIR=. TASK_ID=kysely-window-grouping-helpers
bash harness/smoke_box.sh # smoke kysely on spot m7i.xlarge
bash harness/smoke_box.sh bandit-structured-nosec-directives # different task
DEEP_SWE_DIR=/other/deep-swe bash harness/smoke_box.sh # override source dir
# ~$0.01 EC2, ~3 min wall, self-terminates on EXIT trap, cancels spot request on teardownReceipts land in results/smoke/box/. Expects REWARD 1 from gold patch as the box-infra freeze
gate. See results/smoke/box/RESULT.md for the 2026-05-29 run.
- v1 gold-patch defect audit (2026-05-27):
results/,PREREGISTRATION.md, the oracle ledger, and the confirmed defect list. Four of 113 reference solutions fail their own verifiers. - v1 five-minute codex audit (2026-05-29):
external/codex-audit/— prompt + transcript. - v1.1 receipts (2026-07-07):
results/v1.1/— snapshots of the public v1.1 artifacts (v1-delta.json,leaderboard-live.json,heatmap.json,trials.json.gz) andDERIVATION.md, which re-runs every quantitative claim in the v1.1 write-up.external/codex-audit-v1.1/holds the v1.1 codex prompt + transcript.
Determinacy receipts (a codebase-determinacy sweep of the DeepSWE tasks, per the tool at
kimjune01/determinacy) will land under
results/determinacy/ when that audit runs.