Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions .gitattributes
Original file line number Diff line number Diff line change
Expand Up @@ -24,6 +24,10 @@
*.svg text eol=lf
*.cff text eol=lf

# The fallback selection source is retained byte-for-byte alongside its
# newline-normalized, human-readable JSON projection.
benchmarks/evidence/alignment-external-v1/*.raw binary

# Anything genuinely binary must never be touched.
*.png binary
*.jpg binary
Expand Down
37 changes: 36 additions & 1 deletion CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,40 @@ All notable changes to miniVERL are recorded here. The format follows

## [Unreleased]

## [0.7.0] - 2026-08-10

### External Alignment Gate result

- The first preregistered external-alignment study terminated at starting-
checkpoint selection. Two declared lineages and eight candidates all
measured 0/64 retained JSONNav utility against the unchanged 20% floor. No
checkpoint was selected; teacher qualification, continuation SFT/DPO/KD/OPD
training and the reserved final test did not run. This is a study-design and
precondition finding, not a post-training method comparison.
- Added a schema-valid early-stop result and 512 privacy-safe JSONNav
selection rows. `miniverl pilot --study-result ...` returns
`do_not_continue_this_study` and `insufficient_evidence` without turning one
stopped study into a universal method recommendation.
- Preserved the original fallback selection artifact and published a corrected
lineage-only projection plus correction manifest. The primary and fallback
selection manifests are disclosed as separately generated but byte- and
task-identical, not independent samples.
- Granite Guardian values are explicitly unqualified diagnostics. Granite and
PairRM qualification, PairRM method preference and teacher qualification are
`not_run`; the necessary retained-utility failure does not depend on them.
- Added a generated checkpoint gate matrix and study-flow diagram with mobile
layouts, plus the external-study page, English/Chinese release framing and
browser visual coverage.

### Foundation and artifact hardening

- The bridge validates bundle trees before opening content; refuses symlinks,
reparse points, non-regular or escaping entries and bounded-tree violations;
and distinguishes complete, incomplete and uninspected privacy checks.
- Dataset extension sidecars bind source digest and row count, conversion
revalidates source identity before publication, and row provenance uses
bounded contiguous runs.

### Fixed

- A reward scaffold saved with a UTF-8 byte-order mark is no longer reported as
Expand Down Expand Up @@ -768,7 +802,8 @@ Same-tokenizer only; one trajectory per forward pass; `swap` unavailable for
quantized models; only Qwen3 and Qwen2 architectures tested; single-seed GPU
results. The full list is in `docs/limitations.md`.

[Unreleased]: https://github.com/DaoyuanLi2816/mini-verl/compare/v0.6.3...HEAD
[Unreleased]: https://github.com/DaoyuanLi2816/mini-verl/compare/v0.7.0...HEAD
[0.7.0]: https://github.com/DaoyuanLi2816/mini-verl/compare/v0.6.3...v0.7.0
[0.6.3]: https://github.com/DaoyuanLi2816/mini-verl/compare/v0.6.2...v0.6.3
[0.6.2]: https://github.com/DaoyuanLi2816/mini-verl/compare/v0.6.1...v0.6.2
[0.6.1]: https://github.com/DaoyuanLi2816/mini-verl/compare/v0.6.0...v0.6.1
Expand Down
10 changes: 7 additions & 3 deletions CITATION.cff
Original file line number Diff line number Diff line change
Expand Up @@ -2,8 +2,8 @@ cff-version: 1.2.0
title: "miniVERL: Auditable single-GPU alignment and distillation runtime"
message: "If you use miniVERL in your work, please cite it as below."
type: software
version: 0.6.3
date-released: 2026-08-06
version: 0.7.0
date-released: 2026-08-10
license: Apache-2.0
repository-code: "https://github.com/DaoyuanLi2816/mini-verl"
url: "https://github.com/DaoyuanLi2816/mini-verl"
Expand Down Expand Up @@ -32,7 +32,11 @@ abstract: >-
algorithmic parity with PPO. miniVERL is designed for one personal CUDA GPU,
automatically selects bf16 or fp16, and requires neither Ray nor a cluster.
Published performance is measured on one RTX 4080; other GPU models use the
same code path but remain unmeasured.
same code path but remain unmeasured. The v0.7.0 external-alignment study
terminated at its preregistered checkpoint-selection gate: two declared
lineages and eight candidates all scored 0/64 retained JSONNav utility, so
no teacher qualification, continuation method comparison or reserved final
test was run.
authors:
- family-names: Li
given-names: Daoyuan
Expand Down
67 changes: 49 additions & 18 deletions PROJECT_STATE.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,16 +6,39 @@ and what it printed.

Last updated: 2026-08-10.

Canonical release state: stable `v0.6.3` (`005a4549da713716e64c3ae80ff55fb131519f79`), development `0.7.0.dev0`.
Canonical release state: releasing `v0.7.0`.
Every public version claim is generated from `release-state.yaml` and gated by
`python scripts/release_state.py --check`.

## v0.7.0 External alignment evidence — IN PROGRESS
## v0.7.0 External Alignment Gate — EVIDENCE RELEASE IN PROGRESS

Branch `v0.7.0-foundation`, based on `0bd194600aeb65b90eadd14bfa1ec313aa2a9c36`.
Not yet pushed. Commit author is
`Daoyuan Li <94409450+DaoyuanLi2816@users.noreply.github.com>` on every commit,
matching the 9:1 dominant identity since v0.6.0.
Branch `v0.7.0-evidence-release` starts from the exact post-PR-#52 main commit
`a8272e2b5674e12107461a81d285f1a3d56588a5`. PR #51 merged the public
preregistration/outcome as `c50aa93b95e6fe4a6aa6251491d3c2b5a9480ebe`;
PR #52 corrected the handoff boundary before merge. Every source commit author
is `Daoyuan Li <94409450+DaoyuanLi2816@users.noreply.github.com>`.

Current scientific state is terminal and unchanged: two declared lineages,
eight candidates, 0/64 JSONNav retained utility for every candidate, no
selected checkpoint, no teacher qualification, no continuation method and no
final-test access. The release work is evidence projection and documentation,
not a restarted experiment.

### Evidence-release identities

| evidence | state / SHA-256 |
| --- | --- |
| original fallback selection | preserved byte-for-byte; `53efeb1af196fe8a2fd3733f3f9d6a9ce101fcc76365fc45515adc47cc7d3cd3` |
| corrected fallback projection | lineage metadata only; `6f23de43f03a69275d8bedc9b029a1b728fb2004a9b3a225c66ba1fee671592b` |
| primary / fallback selection manifests | both `e1e165e3547c7784b17e93b7e665df66ea6cafa70bec093a69377bc6683bc20b`; separately generated, task-identical, not independent |
| portable JSONNav selection rows | 512 rows, no prompts/responses/absolute paths; `18d5733e70bfe292e282bd5b6e3fc94869837fab30a151a642aa11c3e4c9d771` |
| early-stop result | schema-validated; digest recorded by the final release validation |
| final test | `not_accessed`, zero tasks scored |

Amendment 4 is explicitly post-selection and pre-release. It changes no
quantitative value, gate, threshold, endpoint or decision. Granite Guardian is
an unqualified diagnostic only; Granite and PairRM qualification, method-level
preference evaluation and teacher qualification are `not_run`.

### Phase A progress

Expand Down Expand Up @@ -209,11 +232,11 @@ harmful compliance 0.37–0.77 with room to fall. Only utility is a hard zero.

### The zero was validated before it was believed

Every rollout ended at `PARSE_ERROR_LIMIT` with **zero tool calls emitted** and
exactly 128 tokens — a uniform deterministic failure across four adapters
including the base model. That signature fits a misconfigured harness as well
as an incapable policy, and 128 = 2 turns x the 64-token per-turn budget is
exactly what a too-small budget would produce.
Every primary-lineage rollout ended at `PARSE_ERROR_LIMIT` with zero valid tool
calls and exactly 128 tokens. Fallback-lineage rollouts emitted two to four
JSONNav tool calls and ended by final answer, max turns or repeated-call limit,
but still solved 0/64 for every adapter. These are distinct observed failure
behaviors, not one uniform signature.

The environment's oracle clears the identical settings 8/8, all
`FINAL_ANSWER`, at 64, 128 and 256 tokens per turn and at both difficulties.
Expand Down Expand Up @@ -257,13 +280,21 @@ Contributor audit over the merged range: 16 commits, sole author
GPU spent across all of v0.7 so far: **1.2 hours** of the 48-hour envelope,
peak reserved 5.25 GiB against the 14.5 GiB gate.

### Next action

Both lineages failed and the outcome is merged. The next session executes
`docs/handoffs/v0.7.0-final-execution-prompt.md`: make `miniverl pilot` return
the recommendation against downstream alignment, write the study page and its
data-bound figures, update both READMEs, add an evidence comment to issue #39
while keeping it open, and release v0.7.0.
### Release action

The evidence-release branch adds the schema-valid early-stop result, portable
selection rows, original/corrected fallback relation, identical-suite
disclosure, `miniverl pilot --study-result`, the study page and data-bound
figures. After its full green validation it will release v0.7.0, add an
evidence comment to issue #39 while keeping the issue open, then advance main
to `0.7.1.dev0` through a separate state-sync PR.

Local release gates on implementation commit `013993a8cd5002a9ed166ba1e6948305f92c1bfb`
plus its quality-record-only update pass 2,098 CPU tests with 6 platform skips
and 21 deselections at 86.35% combined branch coverage. The available RTX 4080
suite passes 8 tests and the opt-in network suite passes 15. Ruff, format,
mypy, actionlint, strict MkDocs, four-viewport browser checks, package/Twine,
clean-install and extracted-sdist gates pass.

No continuation arm, no teacher qualification and no method comparison will
run — none of them is scientifically authorized by the selection outcome.
Expand Down
Loading