Skip to content

Release RecoveryBench v1 - #27

Merged
DaoyuanLi2816 merged 25 commits into
mainfrom
agent/v0.3-recoverybench
Aug 2, 2026
Merged

Release RecoveryBench v1#27
DaoyuanLi2816 merged 25 commits into
mainfrom
agent/v0.3-recoverybench

Conversation

@DaoyuanLi2816

@DaoyuanLi2816 DaoyuanLi2816 commented Aug 2, 2026

Copy link
Copy Markdown
Owner

Summary

  • publish the preregistered three-seed RecoveryBench v1 schema-v3 results, compact task data, paired analysis and generated SVGs
  • add the deterministic six-page technical report and data-bound artifact/hash tests
  • position SFT as competence-building and OPD as a teacher-dependent online mechanism; keep the calculator result as a teacher-qualification case study
  • preserve the legacy calculator JSON and disclose the invalid v1.2 partial run, reverted post-run wall amendment and cycle-capped wall diagnostic

Result

Under eight equal continuation updates, frozen-student KD reached 23.2% strict task success and 22.8% recovery after error; strict fresh-state OPD reached 10.9% and 9.1%. The paired fresh-minus-frozen differences are -12.24 points (95% paired bootstrap -15.89 to -8.59; 384 tasks) and -13.79 points (-20.69 to -6.90; 116 paired error cases).

Fresh OPD averaged 686.80 continuation seconds versus 52.10 seconds for frozen KD. The budget-50 selector queried 49.77% of model-generated positions but did not reduce teacher backbone forwards or wall time. The nominal 50-second artifact is explicitly reported as a cycle-capped diagnostic, not exact equal-time evidence.

Frozen artifacts

  • equal updates: 6ce2e6837e12b99ebc4fad6d27ce3e69c92e295ff3b9b60e0f68c2d308022384
  • equal selected positions: fe4c9afc799724dfe7a32e631676a1e5177c44559a7374d2ea31da135354f137
  • wall diagnostic: 425b0fa568f37b09e61af731d3da5009bd3833bddde6efaf2c66e9dba8355cbe
  • task JSONL: aff96bffc6da27240a852410ac041bd4d95badf34cad030e6f437be1491a55ad
  • paired analysis: 8a6891f74aed80f07ec00d5ea1909895c579346e1abbb1d5d95a354bb46c6b81
  • report PDF: c506300599942445f24b30a4e0d7e01972c75daa7834f6d1eff5b8132dce93af
  • legacy calculator JSON unchanged: 53fc1d4d5b7adee09618d77ad62d4086ba56b78569832d6fc7c3bcd5c2695bbc

Audit scope: 36 task artifacts, 4,608 trajectories and 101,787,618 raw bytes; every embedded digest/size matched, every arm has 128 paired test tasks, frozen checkpoints/datasets are reused exactly, and all 75 fresh updates have matching rollout/current parameter versions.

Validation

  • Ruff check/format, mypy src/miniverl, actionlint and git diff --check
  • 1408 non-GPU/non-network tests, 6 deselected, 86.68% branch coverage
  • 5 GPU tests on RTX 4080; 3 network tests
  • minimum Python 3.10 / training dependency boundary and latest Python 3.13 boundary
  • wheel/sdist build and Twine; clean core and training-extra installs
  • extracted-sdist Ruff/format/mypy plus 1408 tests; rebuilt wheel inventory matches
  • generated artifact byte comparison, Markdown links, JSON parsing, native/README SVG inspection and six-page PDF inspection
  • derived JSON/JSONL generation now pins LF explicitly; committed Git blobs, Linux CI and regenerated Windows bytes share the hashes above

Scope

This is a mechanism study on one Qwen3 pair, one SQLite recovery task family, three seeds and one RTX 4080. It is not an alignment benchmark and does not claim OPD universally loses, offline KD always wins, or local miniVERL execution is equivalent to distributed verl.

@DaoyuanLi2816
DaoyuanLi2816 marked this pull request as ready for review August 2, 2026 23:19
@DaoyuanLi2816
DaoyuanLi2816 merged commit bee82d3 into main Aug 2, 2026
10 checks passed
@DaoyuanLi2816
DaoyuanLi2816 deleted the agent/v0.3-recoverybench branch August 2, 2026 23:19
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant