Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
46 changes: 36 additions & 10 deletions PROJECT_STATE.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ Living build log for **miniVERL** (`mini-verl` / `miniverl` / CLI `miniverl`).
A checkbox is not evidence: every completed item names the command that was run
and what it printed.

Last updated: 2026-08-06.
Last updated: 2026-08-10.

Canonical release state: stable `v0.6.3` (`005a4549da713716e64c3ae80ff55fb131519f79`), development `0.7.0.dev0`.
Every public version claim is generated from `release-state.yaml` and gated by
Expand Down Expand Up @@ -115,9 +115,10 @@ one author.
| training data | **done** (`b0dbaaf`). `Anthropic/hh-rlhf` (MIT), disjoint from every endpoint. 17,416 usable of 20,000 real rows. Over-long examples dropped, never truncated; split from a hash of the example's identity, not its row position |
| SFT candidates | **done** (`08fe30f`). One bounded run, adapters at updates 0/4/8/16, 81 s, peak reserved 4.008 GiB against the 14.5 gate |
| final-test isolation | **done** (`08fe30f`). `prepare_suite(reserved_task_ids=...)` withholds the frozen final-test ids from every selection suite and records the count; an exhausted pool fails closed |
| preregistration | **done** (`d7fbae9`), amended below. Not yet merged, so nothing is frozen and no final-test task has been scored |
| candidate evaluation | **blocked**, see below |
| teacher qualification | not started |
| preregistration | **done**. PR #51 merged as `c50aa93b95e6fe4a6aa6251491d3c2b5a9480ebe`; the early-stop outcome is public and frozen |
| candidate evaluation | **done: checkpoint selection failed**. Both declared lineages completed; every candidate failed the necessary JSONNav retained-utility condition at 0/64 |
| judge qualification | **not run**. Granite Guardian values are unqualified diagnostics only; PairRM qualification and method preference were not run |
| teacher qualification | **not run — scientifically unauthorized** because no starting checkpoint was selected |

### Preregistration amendment 1, and why it exists

Expand Down Expand Up @@ -240,15 +241,40 @@ candidate 0 (the anchor itself, zero continuation updates) also scores 0/64.
swapping it after seeing the result is precisely what preregistration exists to
prevent. This is recorded as a study-design finding for the report.

### Phase C closed

Preregistration PR [#51](https://github.com/DaoyuanLi2816/mini-verl/pull/51)
merged as `c50aa93b95e6fe4a6aa6251491d3c2b5a9480ebe`. That commit is the public
freeze; nothing in the study may move without a further amendment.

Contributor audit over the merged range: 16 commits, sole author
`Daoyuan Li <94409450+DaoyuanLi2816@users.noreply.github.com>`. The two
`Anthropic` matches in commit bodies are the dataset repository name
`Anthropic/hh-rlhf`, not attribution; an attribution-shaped pattern
(`co-authored-by:`, `generated by ai`, `by claude`, `claude code`,
`assisted by`) returns zero hits.

GPU spent across all of v0.7 so far: **1.2 hours** of the 48-hour envelope,
peak reserved 5.25 GiB against the 14.5 GiB gate.

### Next action

Let the fallback lineage finish. If it fails too, publish
`checkpoint_selection_failed` as preregistered, run nothing teacher-dependent
or downstream, and do **not** invent a third lineage. v0.7.0 then publishes the
benchmark infrastructure, the judge qualification, the candidate failure and a
pilot recommendation against downstream alignment.
Both lineages failed and the outcome is merged. The next session executes
`docs/handoffs/v0.7.0-final-execution-prompt.md`: make `miniverl pilot` return
the recommendation against downstream alignment, write the study page and its
data-bound figures, update both READMEs, add an evidence comment to issue #39
while keeping it open, and release v0.7.0.

No continuation arm, no teacher qualification and no method comparison will
run — none of them is scientifically authorized by the selection outcome.

Before release, correct the fallback lineage metadata with original-byte
provenance, disclose that the two separately generated selection suites are
task-identical, and remove every unsupported judge-qualification claim. Add an
evidence comment to issue #39 and keep it open because the method-level and
final-test acceptance gates remain unexecuted.

**No final-test task has been scored.** `first_final_test_access: pending`.
**No final-test task has been scored.** `first_final_test_access: not_accessed`.

## v0.6.3 Security, artifact integrity and release-state hardening

Expand Down
38 changes: 22 additions & 16 deletions docs/handoffs/v0.7.0-final-execution-prompt.md
Original file line number Diff line number Diff line change
Expand Up @@ -20,10 +20,10 @@ Your job is to publish that outcome honestly, release `v0.7.0`, and advance to

```text
stable release: v0.6.3
main: <FINAL_MERGE_SHA>
main: c50aa93b95e6fe4a6aa6251491d3c2b5a9480ebe
development version: 0.7.0.dev0
preregistration merge: <FINAL_MERGE_SHA>
first_final_test_access: pending — and it stays pending
preregistration merge: c50aa93b95e6fe4a6aa6251491d3c2b5a9480ebe
first_final_test_access: not_accessed
```

Audit before acting:
Expand Down Expand Up @@ -108,35 +108,41 @@ it.
- pinned external endpoint governance: four ungated sources at 40-hex revisions,
licences and redistribution decisions recorded, rejected candidates recorded
with reasons;
- four working evaluators — IFEval (independent, 25/25 instruction types, 0 of
- four evaluator adapters — IFEval (independent, 25/25 instruction types, 0 of
834 unscored), XSTest string-match refusal, JBB + Granite Guardian, PairRM;
- IFEval and XSTest executed on the selection split; Granite Guardian was used
only as an unqualified diagnostic; Granite Guardian and PairRM qualification
were not run;
- the frozen 500-task suite with demonstrated final-test disjointness;
- the RewardBench/XSTest overlap exclusion (404 rows);
- three preregistration amendments, each disclosing its timing;
- the checkpoint-selection failure with its full diagnosis;
- `miniverl pilot` recommending against downstream alignment on this evidence.

**Not** published: any method comparison, any teacher qualification, any
continuation-arm result. None of those ran, and the documents must not imply
they did.
**Not** published: any method comparison, judge qualification, teacher
qualification, PairRM method preference or continuation-arm result. None of
those ran, and the documents must not imply they did.

## 5. Remaining work

1. `miniverl pilot` returns `insufficient_evidence` / recommends against
1. Correct the fallback lineage metadata with original-byte provenance, disclose
that the separately generated selection manifests are task-identical, and
label Granite Guardian as an unqualified diagnostic.
2. `miniverl pilot` returns `insufficient_evidence` / recommends against
downstream alignment, citing the selection failure. Add a regression.
2. `docs/alignment-external/alignment-external-v1.md` — the study page. Lead
3. `docs/alignment-external/alignment-external-v1.md` — the study page. Lead
with the question, the design, and the failure. No figure may imply a method
comparison happened.
3. Any figure must be generated from `benchmarks/preregistration/`
4. Any figure must be generated from `benchmarks/preregistration/`
`alignment-external-v1.yaml` and the two selection artifacts. No manual
numbers in SVG. 390px-readable, non-colour encoding, alt text.
4. README and `docs/index.md`: the external study is a *negative* result. Keep
5. README and `docs/index.md`: the external study is a *negative* result. Keep
calculator, RecoveryBench, Consumer Runtime and Alignment Lab as linked case
studies. Do not delete negative historical evidence.
5. English and Chinese together.
6. Close issue #39 with links to the preregistration, the selection artifacts,
the failure page and the release the roadmap item is answered by executed
evidence, even though the answer is "the study could not run as designed".
6. English and Chinese together.
7. Add an evidence comment to issue #39 with links to the preregistration,
selection artifacts, failure page and release. Keep the issue open because
method-level and final-test acceptance gates remain unexecuted.

## 6. Gates

Expand Down Expand Up @@ -168,7 +174,7 @@ private path or fake secret reaches a public artifact.
5. annotated tag `v0.7.0` on the exact validated merge commit;
6. publish via the existing OIDC workflow;
7. verify PyPI hashes, attestations, GitHub Release assets, clean install;
8. close issue #39;
8. comment on issue #39 with the evidence and keep it open;
9. state-sync PR to `0.7.1.dev0` recording the release evidence.

Do not tag if publication authorization or trusted-publisher configuration is
Expand Down