Status: complete. Every bounded V11.1-V11.22 implementation slice is complete. Explicit follow-ons remain owned by their named plans and do not keep the V11 version frame active.
Release target:
- post-
1.6.22planning slice - broad version frame after
1.6.22
Source direction:
- Execution receipt
- Semantic diff and explain
- Doctor finding contract
- JSON output reference
- V10 plan
- V11.1 plan
- V11.2 plan
- V11.3 plan
- V11.4 plan
- V11.5 plan
- V11.6 plan
- V11.7 plan
- V11.8 plan
- V11.9 plan
- V11.10 plan
- V11.11 plan
- V11.12 plan
- V11.13 plan
- V11.14 plan
- V11.15 plan
- V11.16 plan
- V11.17 plan
- V11.18 plan
- V11.19 plan
- V11.20 plan
- V11.21 plan
- V11.22 plan
V11 theme:
- completion truth for agents
- reviewer evidence for agent-authored work
- local, CI, and agent execution convergence
The first concrete V11 implementation slice is:
This slice turns a recurring product signal into an explicit planning surface:
- agents can produce code that looks plausible
- maintainers still reject the PR because the repo never made completion truth explicit enough
The product goal is not "better prompting." The product goal is making the repo say what a correct, safe, complete change actually means.
A large class of rejected agent-authored work is not only a model-quality problem.
It is an execution-governance problem:
- the agent prepared the repo incorrectly
- the agent ran the wrong verification lane
- the repo required services, env, or setup order that were never declared clearly
- the agent crossed unsafe boundaries because the repo did not make them explicit
- the maintainer had to reconstruct whether the change was actually complete
Ota already addresses parts of this, but the current product surface still leaves "what counts as done?" too implicit in many repos.
V11 is the slice for making completion truth and reviewer evidence first-class.
V11 is the broader version frame.
Its first planned implementation slice is V11.1:
- execution governance visibility and proof
Later V11 slices should build on that foundation instead of competing with it.
The next planned V11 slice after that is:
That slice is intentionally the repo-truth convergence layer before the next execution-surface widenings. It should make later work such as container-backed hydration, deterministic bootstrap/materialization, and richer runtime bootstrap ownership proceed from governed evidence instead of ad hoc repo pressure alone.
The next planned runner-enforcement slice after that is:
That slice closes the remaining gap between safe-task and workflow-closure declaration and actual runtime control, so agent-safe execution truth becomes enforceable by the runner instead of staying only a governance and review surface.
The implemented OSS governance slices after that are:
- V11.4: machine-readable governance evaluation output
- V11.6: harness and sandbox capability integration
The following governance continuation is complete for its bounded OSS slice:
- V11.7: audited execution boundary crossings - completed continuation:
canonical semantic scope, the fixed-trust
prebound_filesigned carrier, and the Unix launcher-session broker carrier are implemented. The systemd carrier now provides one-use work-unit authority, selected execution, terminal cleanup, crossing receipt/archive evidence, portable finalization, and immutable Linux/x64 PID 1 pressure for governedrun/upand proof-wide transactions. Its production invocation client, least-privilege protected-history source, independently administered launcher installation, and administrator-driven reboot/fault recovery are pressure-proven. Provider attestation remains optional stronger hardening, and contract-authored crossing declarations are deferred follow-on authoring work. No later slice may treat a crossing record as reusable approval authority.
The completed trust/product follow-ons are:
- V11.3: agent-scoped execution enforcement - implementation and real-repo refusal-canary pressure complete.
- V11.5: CI and merge-gate projection - required lanes, drift, and CI-owned refusal-canary checks complete through the V11.15 GitHub adapter.
- V11.9: governance truth reconciliation and evidence classes
- V11.10: replay-verified baseline trust and last-known-good posture - the scoped first replay carrier and Bedrock pressure target are complete; broader provider-backed replay and baseline-promotion policy remain outside that completed slice.
- V11.11: machine-readable proof boundaries and not-proved scope - the first proof-boundary carrier and Athena seam-control pressure target are complete; proof evidence remains execution-authored and cannot be inherited by a lane that did not execute it.
- V11.12: typed hydration input provenance
- V11.13: generated artifact lineage
- V11.8: sandbox policy compilation from the execution contract - completed at the capability-profile compilation boundary; bounded provider-enforced application was delivered separately in V11.21.
- V11.14: contract-claim assurance - implementation and pressure complete.
The completed provider-adapter slice is:
- V11.15: managed GitHub Actions governance projection - implementation and real-repo pressure complete.
The completed execution-trust slice is:
- V11.16: fresh-boundary setup proof - scoped implementation and declared filesystem-boundary pressure complete.
The completed replay-governance slice is:
- V11.17: trusted replay baseline regeneration - implementation and independent EventCatalog pressure complete, with explicit record, promotion, and detached replay consumption.
The completed lifecycle-proof slice is:
- V11.18: managed lifecycle-sequence proof - implementation, two-manager-family pressure, and independent review complete.
The completed typed-hydration slice is:
- V11.19: typed uv local-project hydration - Dograh proves the nested editable-project, full extras, and ordered dependency-group lane; Marimo independently proves the root-project test-group shape on Linux and macOS while remaining narrowing because it has no lockfile.
The completed policy-governance slice is:
- V11.20: policy-governed replay input identity - implementation and Bedrock/Kylrix pressure complete.
The completed sandbox-enforcement slice is:
- V11.21: enforced sandbox policy application - completed on
1.6.26-implementation. It applies the provider-neutral canonical policy envelope through the bounded stock OCI subset and refuses unsupported authoritative controls before work begins.
The completed OSS authoring slice is:
- V11.22: contract creation and quality UX - completed. It provides deterministic, source-bound contract candidates, fail-closed execution-closure classification, explicit create-new and Git write carriers, and registered lossless upgrades without inventing a second claim model or hosted intelligence dependency.
One deferred harness-integration follow-on is recorded in the completed V11.6 plan: an agent-evaluation disposition profile that prevents preflight policy refusal from being counted as agent capability failure while still allowing an expected refusal to pass an explicitly declared admission-compliance evaluation. It remains inactive until V11.7 closes and real evaluation-harness pressure validates the need; it must derive scoring eligibility from existing governance truth rather than introduce a parallel verdict taxonomy.
That follow-on remains durably owned by V11.6 and inactive pending real harness demand and pressure; it is not part of V11 completion and does not move into V12. Contract-authored crossing requirements are owned by V12.2, while optional provider-attested and non-Linux carrier work is owned by the demand-gated V12.3-V12.5 plans.
Those slices make Ota higher in the stack without abandoning the open execution spec:
- V11.4 publishes portable governance truth
- V11.5 makes CI and merge gates enforce contract-owned completion truth
- V11.6 lets external harnesses enforce Ota’s callable boundary without guessing
- V11.7 makes allowed-but-heavier execution explicit, classifiable, and auditable in OSS before enterprise approval layers build on top
- V11.8 compiles contract-owned execution boundary truth into a provider-neutral capability profile; V11.21 applies the bounded first provider-enforced filesystem and network subset
The completed trust-refinement sequence established:
- V11.9 tightens the trust model so governance fields are emitted from the same decision line that made them, typed by evidence class, decomposed where Ota already knows truthful blocker or gate structure, and checked for post-decision reconciliation instead of drifting into second-read assembled JSON
- V11.10 then strengthens the receipt/baseline trust story so "last known good" means named inputs, exact witness, and replay posture instead of only one historical green outcome
- V11.11 then makes narrow proof honest in machine-readable form so proof artifacts can say what they covered and what they explicitly did not prove
- V11.12 then tightens typed dependency hydration trust where source or feed posture materially changes replayability and execution confidence
- V11.13 names generated source as a contract-owned producer/consumer artifact instead of relying on procedural ordering alone
V11.14 completes the trust-refinement sequence by keeping a maintainer claim, the closure the runner can enforce, observable evidence supporting or contradicting that claim, and the policy decision that admits it separate. Its Athena and Lead Quorum pressure paths prove both supported and unknown assurance outcomes before V11.15 begins.
- explicit completion surfaces for agent-authored change validation
- stronger reviewer-facing evidence for what verification actually ran
- clearer separation between code failure, readiness failure, and contract drift
- better convergence between local, CI, and agent execution truth
- stronger machine-readable stop/review signals for unsafe or incomplete agent outcomes
- do not turn Ota into a generic PR review platform
- do not build hosted human approval workflow as part of this slice
- do not claim that Ota can prove code correctness from receipts alone
- do not collapse execution evidence, semantic diff, and reviewer intent into one structure
- do not make agents autonomous over dangerous tasks just because the verification story improves
Do not frame this as:
- AI agent quality scoring
- agent leaderboard instrumentation
- prompt management
Frame it as:
- completion truth
- reviewer evidence
- execution convergence
The core question is:
- can the repo tell an agent what a correct, safe, complete change looks like?
Today a repo can declare tasks, workflows, readiness, and safe-task boundaries.
What is still weaker than it should be is the explicit answer to:
- what verification lane counts as completion for this class of change
- what evidence must exist before an agent should stop
- what should block "done" even when code changes look locally plausible
V11 should make that surface clearer and more machine-checkable.
Today a maintainer can inspect:
- receipts
- proof artifacts
- doctor output
- semantic diff / snapshot correlation
That is already useful.
What is still weaker than it should be is the direct reviewer answer to:
- what contract/workflow/task the agent believed it was following
- what verification actually ran
- what failed versus what was skipped
- whether the fix changed repo truth, runtime state, or only code
V11 should reduce reviewer reconstruction work further.
Even when the repo declares useful contract truth, maintainers can still end up with:
- one local path
- one CI path
- one agent path
V11 should keep pushing toward one explicit execution contract instead of three partially aligned conventions.
The repo should be able to say more clearly:
- this task is safe
- this effect is external or destructive
- this workflow is verification, not setup
- this repo is not ready, so code-level completion claims should stop here
V11 should strengthen those stop/review semantics.
Define a clearer contract-owned surface for completion truth.
This should let a repo say:
- which task or workflow is the canonical completion lane
- which verification steps must pass after changes
- when "done" is not satisfied even if the code change itself compiles or tests partially
Design bar:
- completion truth should be repo-owned
- completion truth should be machine-readable
- completion truth should reuse existing task/workflow structure where possible
- completion truth should not require reviewers to infer intent from prose alone
Widen receipts / proof / JSON summaries so reviewers can answer:
- what path the agent selected
- what actually executed
- what was skipped
- what contract snapshot and semantic assumptions were in play
- whether the failure or stop condition was code, setup, readiness, or contract drift
This is not a new generic reporting engine. It is the next trust layer on top of existing receipt and proof surfaces.
Push repos harder toward one canonical verification truth across:
- local execution
- CI workflows
- agent execution
Expected direction:
- reuse repo-owned task/workflow truth
- reduce handwritten duplication in CI
- warn more clearly when the repo's public or machine-facing execution stories split
Improve machine-readable and human-readable signals for:
- incomplete verification
- readiness blockers
- skipped required proof
- unsafe mutation boundaries
- task paths that should not be treated as autonomous completion
This should sharpen the line between:
- change attempted
- change verified
- repo ready
- safe to merge
V11 should make Ota better at answering these directly:
- What does this repo require before a change can be considered complete?
- What exact verification path did the agent run?
- Did the repo become ready, or did the agent stop in an unready state?
- Did the fix change code, contract truth, or only local runtime state?
- Is the current PR failure a code problem, a readiness problem, or contract drift?
- Is this agent outcome reviewable as a safe completion candidate, or should it have stopped earlier?
- Define the completion-truth contract surface.
- Publish machine-readable reviewer evidence for selected path and outcome.
- Tighten stop/review semantics around incomplete verification and unready repo state.
- Add governance for local/CI/agent execution drift.
- Pressure-test on real repos with agent-facing verification lanes.
This order keeps the repo-owned truth first, the evidence second, and the stricter governance last.
- a repo can declare a canonical completion lane without relying on prose alone
- Ota can publish what verification actually ran in a reviewer-useful way
- maintainers can distinguish code failure from readiness failure from contract drift more directly
- local, CI, and agent execution truth have a clearer contract-bound convergence path
- incomplete or unsafe agent outcomes surface stop/review signals earlier and more honestly