Skip to content

Latest commit

 

History

History
433 lines (309 loc) · 17.9 KB

File metadata and controls

433 lines (309 loc) · 17.9 KB

V11 Plan

Status: complete. Every bounded V11.1-V11.22 implementation slice is complete. Explicit follow-ons remain owned by their named plans and do not keep the V11 version frame active.

Release target:

  • post-1.6.22 planning slice
  • broad version frame after 1.6.22

Source direction:

V11 theme:

  • completion truth for agents
  • reviewer evidence for agent-authored work
  • local, CI, and agent execution convergence

The first concrete V11 implementation slice is:

This slice turns a recurring product signal into an explicit planning surface:

  • agents can produce code that looks plausible
  • maintainers still reject the PR because the repo never made completion truth explicit enough

The product goal is not "better prompting." The product goal is making the repo say what a correct, safe, complete change actually means.

Problem statement

A large class of rejected agent-authored work is not only a model-quality problem.

It is an execution-governance problem:

  • the agent prepared the repo incorrectly
  • the agent ran the wrong verification lane
  • the repo required services, env, or setup order that were never declared clearly
  • the agent crossed unsafe boundaries because the repo did not make them explicit
  • the maintainer had to reconstruct whether the change was actually complete

Ota already addresses parts of this, but the current product surface still leaves "what counts as done?" too implicit in many repos.

V11 is the slice for making completion truth and reviewer evidence first-class.

Version structure

V11 is the broader version frame.

Its first planned implementation slice is V11.1:

  • execution governance visibility and proof

Later V11 slices should build on that foundation instead of competing with it.

The next planned V11 slice after that is:

That slice is intentionally the repo-truth convergence layer before the next execution-surface widenings. It should make later work such as container-backed hydration, deterministic bootstrap/materialization, and richer runtime bootstrap ownership proceed from governed evidence instead of ad hoc repo pressure alone.

The next planned runner-enforcement slice after that is:

That slice closes the remaining gap between safe-task and workflow-closure declaration and actual runtime control, so agent-safe execution truth becomes enforceable by the runner instead of staying only a governance and review surface.

The implemented OSS governance slices after that are:

The following governance continuation is complete for its bounded OSS slice:

  • V11.7: audited execution boundary crossings - completed continuation: canonical semantic scope, the fixed-trust prebound_file signed carrier, and the Unix launcher-session broker carrier are implemented. The systemd carrier now provides one-use work-unit authority, selected execution, terminal cleanup, crossing receipt/archive evidence, portable finalization, and immutable Linux/x64 PID 1 pressure for governed run/up and proof-wide transactions. Its production invocation client, least-privilege protected-history source, independently administered launcher installation, and administrator-driven reboot/fault recovery are pressure-proven. Provider attestation remains optional stronger hardening, and contract-authored crossing declarations are deferred follow-on authoring work. No later slice may treat a crossing record as reusable approval authority.

The completed trust/product follow-ons are:

The completed provider-adapter slice is:

The completed execution-trust slice is:

The completed replay-governance slice is:

The completed lifecycle-proof slice is:

The completed typed-hydration slice is:

  • V11.19: typed uv local-project hydration - Dograh proves the nested editable-project, full extras, and ordered dependency-group lane; Marimo independently proves the root-project test-group shape on Linux and macOS while remaining narrowing because it has no lockfile.

The completed policy-governance slice is:

The completed sandbox-enforcement slice is:

  • V11.21: enforced sandbox policy application - completed on 1.6.26-implementation. It applies the provider-neutral canonical policy envelope through the bounded stock OCI subset and refuses unsupported authoritative controls before work begins.

The completed OSS authoring slice is:

  • V11.22: contract creation and quality UX - completed. It provides deterministic, source-bound contract candidates, fail-closed execution-closure classification, explicit create-new and Git write carriers, and registered lossless upgrades without inventing a second claim model or hosted intelligence dependency.

One deferred harness-integration follow-on is recorded in the completed V11.6 plan: an agent-evaluation disposition profile that prevents preflight policy refusal from being counted as agent capability failure while still allowing an expected refusal to pass an explicitly declared admission-compliance evaluation. It remains inactive until V11.7 closes and real evaluation-harness pressure validates the need; it must derive scoring eligibility from existing governance truth rather than introduce a parallel verdict taxonomy.

That follow-on remains durably owned by V11.6 and inactive pending real harness demand and pressure; it is not part of V11 completion and does not move into V12. Contract-authored crossing requirements are owned by V12.2, while optional provider-attested and non-Linux carrier work is owned by the demand-gated V12.3-V12.5 plans.

Those slices make Ota higher in the stack without abandoning the open execution spec:

  • V11.4 publishes portable governance truth
  • V11.5 makes CI and merge gates enforce contract-owned completion truth
  • V11.6 lets external harnesses enforce Ota’s callable boundary without guessing
  • V11.7 makes allowed-but-heavier execution explicit, classifiable, and auditable in OSS before enterprise approval layers build on top
  • V11.8 compiles contract-owned execution boundary truth into a provider-neutral capability profile; V11.21 applies the bounded first provider-enforced filesystem and network subset

The completed trust-refinement sequence established:

  • V11.9 tightens the trust model so governance fields are emitted from the same decision line that made them, typed by evidence class, decomposed where Ota already knows truthful blocker or gate structure, and checked for post-decision reconciliation instead of drifting into second-read assembled JSON
  • V11.10 then strengthens the receipt/baseline trust story so "last known good" means named inputs, exact witness, and replay posture instead of only one historical green outcome
  • V11.11 then makes narrow proof honest in machine-readable form so proof artifacts can say what they covered and what they explicitly did not prove
  • V11.12 then tightens typed dependency hydration trust where source or feed posture materially changes replayability and execution confidence
  • V11.13 names generated source as a contract-owned producer/consumer artifact instead of relying on procedural ordering alone

V11.14 completes the trust-refinement sequence by keeping a maintainer claim, the closure the runner can enforce, observable evidence supporting or contradicting that claim, and the policy decision that admits it separate. Its Athena and Lead Quorum pressure paths prove both supported and unknown assurance outcomes before V11.15 begins.

Included capabilities

  • explicit completion surfaces for agent-authored change validation
  • stronger reviewer-facing evidence for what verification actually ran
  • clearer separation between code failure, readiness failure, and contract drift
  • better convergence between local, CI, and agent execution truth
  • stronger machine-readable stop/review signals for unsafe or incomplete agent outcomes

Non-goals

  • do not turn Ota into a generic PR review platform
  • do not build hosted human approval workflow as part of this slice
  • do not claim that Ota can prove code correctness from receipts alone
  • do not collapse execution evidence, semantic diff, and reviewer intent into one structure
  • do not make agents autonomous over dangerous tasks just because the verification story improves

Product framing

Do not frame this as:

  • AI agent quality scoring
  • agent leaderboard instrumentation
  • prompt management

Frame it as:

  • completion truth
  • reviewer evidence
  • execution convergence

The core question is:

  • can the repo tell an agent what a correct, safe, complete change looks like?

Core product gaps

1. Completion truth is still too implicit

Today a repo can declare tasks, workflows, readiness, and safe-task boundaries.

What is still weaker than it should be is the explicit answer to:

  • what verification lane counts as completion for this class of change
  • what evidence must exist before an agent should stop
  • what should block "done" even when code changes look locally plausible

V11 should make that surface clearer and more machine-checkable.

2. Reviewer evidence is still too reconstructive

Today a maintainer can inspect:

  • receipts
  • proof artifacts
  • doctor output
  • semantic diff / snapshot correlation

That is already useful.

What is still weaker than it should be is the direct reviewer answer to:

  • what contract/workflow/task the agent believed it was following
  • what verification actually ran
  • what failed versus what was skipped
  • whether the fix changed repo truth, runtime state, or only code

V11 should reduce reviewer reconstruction work further.

3. Local, CI, and agent truth still drift too easily

Even when the repo declares useful contract truth, maintainers can still end up with:

  • one local path
  • one CI path
  • one agent path

V11 should keep pushing toward one explicit execution contract instead of three partially aligned conventions.

4. Unsafe or incomplete agent stopping conditions are still too soft

The repo should be able to say more clearly:

  • this task is safe
  • this effect is external or destructive
  • this workflow is verification, not setup
  • this repo is not ready, so code-level completion claims should stop here

V11 should strengthen those stop/review semantics.

Proposed execution slices

1. Completion contract surface

Define a clearer contract-owned surface for completion truth.

This should let a repo say:

  • which task or workflow is the canonical completion lane
  • which verification steps must pass after changes
  • when "done" is not satisfied even if the code change itself compiles or tests partially

Design bar:

  • completion truth should be repo-owned
  • completion truth should be machine-readable
  • completion truth should reuse existing task/workflow structure where possible
  • completion truth should not require reviewers to infer intent from prose alone

2. Reviewer evidence surface

Widen receipts / proof / JSON summaries so reviewers can answer:

  • what path the agent selected
  • what actually executed
  • what was skipped
  • what contract snapshot and semantic assumptions were in play
  • whether the failure or stop condition was code, setup, readiness, or contract drift

This is not a new generic reporting engine. It is the next trust layer on top of existing receipt and proof surfaces.

3. Execution convergence governance

Push repos harder toward one canonical verification truth across:

  • local execution
  • CI workflows
  • agent execution

Expected direction:

  • reuse repo-owned task/workflow truth
  • reduce handwritten duplication in CI
  • warn more clearly when the repo's public or machine-facing execution stories split

4. Stronger stop/review semantics

Improve machine-readable and human-readable signals for:

  • incomplete verification
  • readiness blockers
  • skipped required proof
  • unsafe mutation boundaries
  • task paths that should not be treated as autonomous completion

This should sharpen the line between:

  • change attempted
  • change verified
  • repo ready
  • safe to merge

Proposed operator questions

V11 should make Ota better at answering these directly:

  • What does this repo require before a change can be considered complete?
  • What exact verification path did the agent run?
  • Did the repo become ready, or did the agent stop in an unready state?
  • Did the fix change code, contract truth, or only local runtime state?
  • Is the current PR failure a code problem, a readiness problem, or contract drift?
  • Is this agent outcome reviewable as a safe completion candidate, or should it have stopped earlier?

Rollout order

  1. Define the completion-truth contract surface.
  2. Publish machine-readable reviewer evidence for selected path and outcome.
  3. Tighten stop/review semantics around incomplete verification and unready repo state.
  4. Add governance for local/CI/agent execution drift.
  5. Pressure-test on real repos with agent-facing verification lanes.

This order keeps the repo-owned truth first, the evidence second, and the stricter governance last.

Acceptance bar

  • a repo can declare a canonical completion lane without relying on prose alone
  • Ota can publish what verification actually ran in a reviewer-useful way
  • maintainers can distinguish code failure from readiness failure from contract drift more directly
  • local, CI, and agent execution truth have a clearer contract-bound convergence path
  • incomplete or unsafe agent outcomes surface stop/review signals earlier and more honestly