Repository navigation
release: v2026.9.19 - #8436
Merged
Merged
release: v2026.9.19#8436
Conversation
houko
enabled auto-merge (squash)
September 19, 2026 17:45
houko
force-pushed
the
chore/bump-version-2026.9.19
branch
from
September 20, 2026 12:20
f7de9c9 to
1e93e25
Compare
CHANGELOG.md carried a literal (#PR) placeholder for the ChunkErrorBoundary fix, sourced from changelog.d/fixed/8406-router-error-boundary-props.md. Fill in the actual PR number (#8406) so the release notes don't ship a broken reference.
…ed entry The release job skips Highlights generation when the `claude` CLI is unavailable on the runner, which is what happened here, so the 2026.9.19 section shipped without the summary every previous release opens with. #8406 appeared twice because its fragment ended with the literal `(#PR)` placeholder rather than `(#8406)`, so `curated_pr_refs` could not match it and the generated PR-title line was emitted alongside the curated bullet. The placeholder is already fixed on this branch; this drops the now-redundant generated line.
houko
force-pushed
the
chore/bump-version-2026.9.19
branch
from
September 20, 2026 12:22
1e93e25 to
37f50ac
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Release v2026.9.19
56 PRs from 3 contributors since v2026.9.14.
Added
agent_spawngains aprofileparameter that pins the spawned agent onto a named model profile.This is the case model profiles were built for: the goal loop auto-spawns helpers, and without it a verifier whose only job is to answer "is this done?" inherits its parent's expensive model.
Passing
profile = "quick"now writes that profile's provider and model onto the child's manifest instead of letting it inherit the default.Naming a profile that does not exist fails with the list of profiles that do, rather than spawning an agent on the wrong model.
The parameter is honoured whether or not
[model_router] enabledis set: that switch governs the automatic per-turn router, while naming a profile here is an explicit choice, so a cheap sub-agent does not require turning on routing for every agent (feat: complexity-based model routing (ModelRouter, model profiles, per-agent router overrides) #7757) (@DaBlitzStein)The common deployment reality is a cheap model that handles most turns and an expensive one that should only see the hard ones; until now that choice was static per agent, so operators either overpaid on every trivial turn or underserved the hard ones.
A
ModelProfilebinds task tags to a provider/model pair, a cost tier and a complexity ceiling; the kernel scores each turn from its text and picks the best profile the agent is permitted to use.Builtin profiles ship as an asset and are overridden per-name from
~/.librefang/model_profiles.toml, which is re-read when its mtime moves so an edit needs no restart.Off by default: set
[model_router] enabled = trueinconfig.tomlandmode = "flexible"in an agent's[model]block.Per-agent constraints live in
[model.router_override]—allowed_profileslimits the choice,cost_budgetcaps the tier,default_profileis the fallback, andfixed = trueopts the agent out entirely.A fallback profile is re-checked against those same constraints, so a
default_profilecan never spend past an agent's cost budget.Configurable from all four surfaces:
GET/PUT /api/agents/{id}/model_routingandGET /api/model-router/profileson the API, a Routing tab on the dashboard's agent detail, anreditor on the TUI agent detail screen, andlibrefang agent routing/routing-set/routing-profileson the CLI (feat: complexity-based model routing (ModelRouter, model profiles, per-agent router overrides) #7757) (@DaBlitzStein)A run ended the moment the agent wrote
GOAL_DONE, which left the worker as the sole judge of its own work — the one check a long-horizon loop most needs, and the one it did not have.Setting
loop_engineeringon a goal adds two judges that are not the worker.A verifier agent (
verify_agent_id) reads each iteration's output and returnsVERDICT: PASS|FAIL|NEEDS_REWORK; a rejection goes back to the generator carrying the verifier's stated reason, up toverify_max_retriesrework rounds, and until the verifier passes the workGOAL_DONEdoes not end the run.An evaluator model (
evaluator_model) makes one cheap yes/no read of the goal against the latest output, and can conclude the goal is met even when the agent never claimed it.Because the verifier now stands between the agent and the end of the run,
GOAL_DONEchanges meaning: it is a request to finish, granted once the verifier passes the work it is attached to, rather than the finish itself.The agent can also record a reusable lesson with
GOAL_LEARNED: <one sentence>; captured lessons are replayed into later iterations' prompts, persisted to the shared store, and filed as a draft skill in the workshop'spending/queue.Turning a run's lessons into a skill an agent actually loads still takes a human
librefang skill pending approve— an autonomous loop proposes, it does not grant itself authorship.All three are inert unless the goal opts in, so a goal that does not ask for them sends the same prompt and makes the same single LLM call per iteration it always did.
Sub-agents are delegated rather than provisioned: the prompt directs the agent to its own
agent_spawn/agent_sendtools, which run under the capabilities its operator granted it, and neither the runner nor the API ever creates an agent on a caller's behalf.First of the five PRs splitting the closed feat(goals): autonomous goal runner with loop engineering, /goal CLI, TUI, and channels #6505 (feat(goals): gate autonomous goal runs on a verifier and an evaluator #7785) (@DaBlitzStein)
POST /api/goals/{id}/pauseandPOST /api/goals/{id}/resume.Until now
stopwas the only way to halt a run and it discarded the run's state, so an operator who wanted to free an agent for an hour could only express that as "throw away the last forty iterations and start over".A pause lets the loop finish the turn it is on, then checkpoints its iteration count, progress and start time, so the run stays visible in
GET /api/goals/{id}/runin the newpausedphase while nothing is being spent on it, and a resume continues the same run under the same iteration budget rather than beginning a new one (feat(goals): add pause/resume and a configurable loop cadence #7973) (@DaBlitzStein)tick_interval_secson the goal itself, accepted byPOST /api/goalsandPUT /api/goals/{id}.Every goal previously ticked at the same hard-wired cadence, which is the wrong rate at both ends: a goal whose progress depends on something outside the daemon burned turns re-reading an unchanged world, and one that needs to react quickly could not be told to.
An out-of-range value is refused rather than clamped, so an operator who types a cadence learns their number was rejected instead of discovering later that the loop runs at a rate they never chose (feat(goals): add pause/resume and a configurable loop cadence #7973) (@DaBlitzStein)
The page now polls a selected run every 3 seconds while it is active — including while paused at an operator gate — and stops once it reaches a terminal state, so an operator watching a run does not need to reload to see it advance past an approval or finish. (feat(dashboard): workflow run timeline and task board improvements #7997) (@DaBlitzStein)
The controls call the new
POST /api/goals/{id}/pauseand/api/goals/{id}/resumeendpoints, pause stays a run-level phase rather than a goal status, and the badge and progress display keep deriving from the run phase (feat(dashboard): add pause/resume controls to the goals page #8029) (@DaBlitzStein)[skills.promotion]config section for the registry promotion flow, which previously had almost every GitHub-side value hardcoded or derived at runtime.api_base_urlmakes "Propose to Registry" usable on GitHub Enterprise at all, where a compiled-inapi.github.comleft it with no workaround.commit_author_name/commit_author_emailfix the attribution of the pushed commits, which GitHub otherwise credits to whoever owns the token — wrong for a shared or service token, and not something an operator could correct.fork_owner,base_branch,head_branch_prefixand amodeofforkordirect_pushcover the organisation-owned fork, the non-default target branch, the branch-naming convention and the internal registry nobody is meant to fork.Every field is optional and reproduces the previous behaviour when unset, so an installation that configures nothing sees no change. (feat(skills): make the registry promotion GitHub target configurable, and align the CLI's token resolution #8179, Registry promotion: GitHub target, branch, commit author and API base URL are hardcoded or derived, not configurable #8163) (@DaBlitzStein)
p, which it could already show but not act on.mainalready colours a paused run yellow and shipstui-goals-phase-pausedandtui-goals-run-paused, so before this an operator watching from a terminal saw a run somebody had paused from the dashboard, correctly labelled, with no key that touched it.The key reads the live run phase rather than the goal document: a running goal pauses, a paused one resumes, and one doing neither is left alone rather than being started, because
sis the key that starts a run and a pause key that quietly launched one would be a surprise on a screen where "stopped" and "paused" sit next to each other.Resume sends no body, which is the daemon's "keep the cap the paused run was already under" path — re-budgeting a resumed run is a deliberate act and belongs to a surface that can ask for the number, not to a single keypress.
(feat(tui): pause and resume a goal run from the Goals screen #8224) (@DaBlitzStein)
A three-line answer used to take most of a laptop viewport: 24 px of padding on every side of a scrolling column, and type sized for a phone.
The padding and the header type come down, and a control in the conversation header scales the whole transcript between 75% and 125%, remembered between sessions.
Readable type is a per-person, per-display setting — the size that reads well on a 27" panel wastes a 13" one — so the point is the control rather than the new default; a better default only moves whose display it is wrong for.
The scale reaches the transcript only, so making the text smaller never makes the composer or the controls harder to hit. (feat(dashboard): make the chat transcript's size adjustable, and stop spending the space #8299) (@DaBlitzStein)
The
uniongit merge driver concatenates text instead of detecting conflicts, so two branches that each add a key to the same.ftlfile merge cleanly and leave both copies behind.Fluent rejects the whole resource for that language when that happens: a duplicate in the default English pack panics
i18n::initand takes down every test that touches it, while a duplicate in another pack fails silently and falls back to English for that entire language.Both failure modes have already happened once each, discovered only after the fact (test(cli): guard against duplicate Fluent message keys in locale files #8353) (@DaBlitzStein)
GET /api/agentsandGET /api/agents/{id}now carryprovisioned, naming the deployment file that declares the agent, ornullwhen it is the operator's own.Eleven manifest-writing routes answer
423 Lockedon a provisioned agent and the kernel has always known which agents those are, but the payload never said, so a client could not tell before trying: an operator would type an emoji and save, or pick an image and upload the whole thing, only to be refused at the end by something that was never going to work.sourceis the declaring file, which is the one thing a surface needs beyond "you cannot" — it says where to go and change it instead.The field is present and
nullrather than absent when provisioning is switched off, so a client never has to distinguish two spellings of the same answer.The ten of those eleven routes whose OpenAPI said nothing about the refusal now document the
423, so a generated client stops treating it as an unmodelled failure (feat(api): carry an agent's provisioning provenance in its payload #8379) (@houko)SCHEMA_VERSION.A duplicate number used to be invisible:
run_step!fires oncurrent_version < N, so a second step claiming a takenNnever runs on an installation already past it, while a fresh database — which is the only kind CI ever creates — runs every step and shows nothing wrong.The feature would then return 500 on upgraded installations and work perfectly on new ones.
The convention the guard enforces is written down in
docs/development/database-migrations.md, along with why the existing gap check reports the symptom under a message about audit rows.(test(memory): fail the build on a duplicated, skipped or misnamed migration step #8355) (@DaBlitzStein)
Fixed
context_window,max_output_tokensor[model.extra_params]on an agent that inherits the global model.The boot restore shares
clear_stale_provider_overrideswith the model picker and the router, but called it under a branch whose first condition is true for a row already on thedefaultsentinel — where the two assignments above it restate what is there and no endpoint moves.Clearing there is not repointing hygiene: the branch runs on every boot, so a deliberate override disappeared on the next restart and the following
save_agentmade the loss permanent, with the dashboard showing the field blank as if the save had failed.The call is now gated on an actual repoint, which is what both sibling call sites already do (feat(kernel): complexity-based model routing with per-agent profile overrides #7781) (@DaBlitzStein)
set_agent_model— now drops the previous endpoint'scontext_windowandmax_output_tokensalong with its credentials, instead of carrying them onto the new provider.The credential half has been cleared on a provider switch since style: apply cargo fmt to runtime drivers #2380; the limits were not, so moving an agent from a large-window endpoint to a smaller one left the old window attached to the new model and every turn was built against a capacity the new provider does not have.
Both halves are one shared list now, the same one the model router uses and the same one the boot-time restore of a legacy agent to the
defaultprovider uses, so a limit added later cannot be cleared on one path and forgotten on the other two.That list also drops
extra_params, which is flattened verbatim into the request body and only means anything to the provider it was set for — carrying a key like Qwen'senable_memoryonto Anthropic sends a parameter it rejects rather than ignores.A model swap within the same provider still leaves both alone: on one endpoint they are a deliberate per-agent override, not a leftover (feat(kernel): complexity-based model routing with per-agent profile overrides #7781) (@DaBlitzStein)
max_output_tokensnow clears the previous model's output cap the same way it already clearedcontext_window, instead of leaving a routed model capped at the endpoint it just switched away from.PUT /api/agents/{id}/model_routingnow 404s for an agent id that does not exist or that the caller cannot see, matching the GET side of the same endpoint, instead of falling through to a 400 that reads like a malformed request body.The TUI's embedded (in-process) model routing editor no longer wipes an agent's
default_profilefallback on every save — it was never loaded into the editor to begin with, so any edit through an embedded TUI silently cleared it.The same editor also no longer wipes
fixed, the per-agent opt-out that keeps the router from touching an agent at all — every save through an embedded TUI was silently re-enabling routing on an agent an operator had explicitly excluded from it.The tier router (the older
[routing]/[default_routing]model-selection path, separate from the profile router above) now also drops the previous provider'sapi_key_env,base_url,context_windowandmax_output_tokenswhen it switches an agent to a model hosted by a different provider, instead of sending the routed request out under the old provider's credentials and endpoint limits (feat(kernel): complexity-based model routing with per-agent profile overrides #7781) (@DaBlitzStein)librefang agent routing-set --json/routing-show, now surface an agent'sfixedrouter opt-out — it was readable through the API but invisible on every client, so a routing panel showing a fully configured allowlist and budget could still route nothing and nowhere would say why.Routing a profile onto the exact provider and model an agent already has no longer clears a manually-set
context_windowormax_output_tokensfor an endpoint that did not change; only a profile that actually switches the model, or one that supplies its own limit, updates them now.librefang agent routing-setno longer clears the cost budget when only--profilesis passed, or the allowlist when only--budgetis passed — each flag now writes only the field it names, the same "don't touch what you didn't edit" contract the PUT endpoint and the TUI already follow; passing an explicit empty string is still how each one is cleared on purpose.A profile tag differing only in case from another (
"Review"next to"review") no longer double-counts as two tag hits and outranks a correctly written, higher-priority profile that matched the same single word — tags are normalized once when the profile catalog loads instead of on every comparison.The model-routing catalog gate no longer treats "this provider's models haven't synced yet" the same as "nobody has ever described this provider" — a declared-but-unsynced provider (the fresh-offline-install case the gate exists for) still declines an unresolvable model id instead of letting it through, and a model id shared with a different provider's catalog entry no longer resolves as a match for the wrong one (feat(kernel): complexity-based model routing with per-agent profile overrides #7781) (@DaBlitzStein)
Both matchers tested the lowercased task text with
contains, which meantping— a simple-complexity keyword and a tag on the cheapquickprofile — fired on helping, shipping, mapping, grouping and developing, whileformatfired on information,liston specialist, andcounton account.The effect was not a near miss: "Investigate why grouping fails" lost 0.05 of complexity to the phantom
pinghit, scored below the threshold, was capped at the cheapest tier, and then matchedquickon that same phantom hit — an investigation task sent to the cheapest model and reported in the logs and the routing UI as a deliberate tag match.Matching on words routes it through the explicit fallback chain instead, so with no
default_profileconfigured the agent simply keeps its own model.Tasks whose wording only ever contained a keyword inside a longer word will now score differently and may select a different profile; that is the point, but check
default_profileis set if you were relying on the accidental matches (feat(kernel): complexity-based model routing with per-agent profile overrides #7781) (@DaBlitzStein)evaluator_modelno longer stores an empty string where updating the same goal would have removed the field.Update treats a blank value as the signal to clear the key, and the goal runner filters it again on read, so nothing downstream misbehaved — but a created goal round-tripped differently from an updated one and
GET /api/goalshanded the dashboard a value no update would ever have written.Create now drops a blank one the way its siblings do (feat(goals): gate autonomous goal runs on a verifier and an evaluator #7785) (@DaBlitzStein)
loop_engineeringgoal graded an empty statement for a title-only goal, since the dashboard's create form leaves the description optional and the evaluator prompt only ever carried the description.One plausible-looking iteration was enough to draw a YES with nothing to judge it against, closing the goal on iteration one.
The judge now falls back to the title when the description is empty, matching the worker's own prompt, and skips the call entirely when both are empty rather than asking a model to grade nothing (feat(goals): gate autonomous goal runs on a verifier and an evaluator #7785) (@DaBlitzStein)
loop_engineeringgoal is no longer reported as a missing one, and a rework turn that fails every time no longer burns the whole iteration budget in silence.Verifier dispatch failures were all funnelled into a single "unreachable" counter without ever being classified, so a verifier whose provider was merely rate-limiting — a different provider, key or quota bucket than the generator's — ended the run as
Stoppedwithverifier unreachableonlast_errorafter five iterations that each paid for a full generator turn first.The operator was sent looking for a deleted agent, and
RateLimited, the phase the dashboard renders as the retry-later signal, never fired for that leg.Rate limits now keep their own shorter streak there, matching what the generator leg has always done.
The rework turn was the third leg and had no breaker at all: a run of "generator turn fine, verdict FAIL, rework dispatch fails" repeated to the iteration cap and reported the cap as the reason it stopped, with no cause attributed.
It does not take an exotic failure to reach — the rework prompt goes into a session one turn longer than the opening one that just succeeded, so a context-length limit surfaces there first and then repeats every iteration.
Each leg counts its own consecutive failures, because a healthy dispatch on one leg is no evidence at all about the other (feat(goals): gate autonomous goal runs on a verifier and an evaluator #7785) (@DaBlitzStein)
Captured
GOAL_LEARNED:lessons were appended before the verifier ran and again for every reworked reply, and nothing downstream collapsed them — the skill workshop's own de-duplication compares whole candidates, not the entries within one.An iteration that recorded a lesson, drew a
FAIL, and repeated the lesson in its corrected reply therefore stored it twice or more, and those copies crowded genuinely distinct earlier lessons out of the small window replayed into later prompts, out of the stored document, and out of the numbered list a human reads before approving the draft skill.A reworked reply already replaces the rejected one everywhere else, so its lessons now replace them too, and the run refuses a lesson whose text it already holds.
Lessons from an iteration the verifier rejected are still kept, deliberately: unlike
GOAL_DONEthey close nothing and make no claim about the work being graded, and what an attempt that did not land taught is exactly what a human reviewing the draft wants to see (feat(goals): gate autonomous goal runs on a verifier and an evaluator #7785) (@DaBlitzStein)GOAL_LEARNED:lessons, because they were stored under a key scoped to the goal rather than to the run.Lessons are now keyed by the run's own start time as well as the goal id, so re-running a goal no longer costs the operator lessons they may never have read (feat(goals): gate autonomous goal runs on a verifier and an evaluator #7785) (@DaBlitzStein)
GoalRunner::start's return value and hardcodedtrue.A goal that vanished between the API handler's read and the runner's own load — a delete racing a start — made
/goalfrom the dashboard chat, a channel bridge or the TUI print "Goal created and started" for a run that never started, and left the API's ownstartedcheck on that path permanently dead.The refusal now reaches the caller (feat(goals): gate autonomous goal runs on a verifier and an evaluator #7785) (@DaBlitzStein)
Every failing tick was retried until the iteration cap, so a run pointed at an agent that had been removed, or at a provider whose key had been revoked, spent its full budget failing identically each time and then reported the cap as the reason it stopped.
Five consecutive failures that are not rate limits now end the run in
Stopped, with the underlying error left on the run'slast_error— the operator gets the fault instead of a healthy-looking exhausted budget.Rate limits keep their own separate streak, so a provider throttling a run still ends it as
RateLimited(feat(goals): gate autonomous goal runs on a verifier and an evaluator #7785) (@DaBlitzStein)The dashboard's verifier picker offered the goal's own assigned agent and neither goal endpoint compared the two ids, so the pair could be saved and looked entirely healthy afterwards.
With both ids equal the verdict prompt lands in the same persistent session that produced the work one turn earlier, so the agent grades itself with its own output still in context and
VERDICT: PASSis the expected answer — while the run API reports a configured verifier and the dashboard shows the loop-engineering badge, giving the operator positive confirmation of a check that is not checking.Every iteration also left the verification exchange and its own verdict in the worker's history for later iterations to build on.
Both endpoints now reject the pair, update comparing the ids the write would actually leave on the goal rather than only the ones in the payload, and the picker no longer offers the assigned agent in the first place (feat(goals): gate autonomous goal runs on a verifier and an evaluator #7785) (@DaBlitzStein)
verify_max_retriesandverify_agent_idon the goal endpoints skipped the boundary validation their siblings already had.A
verify_max_retriesaboveu32::MAXsilently truncated to a small number instead of being rejected, and a negative or fractional value was indistinguishable from an absent field, unlikemax_iterations, which already rejected both.A non-string
verify_agent_idwas silently dropped instead of rejected, and update's hand-rolled check had the same gap, unlikeagent_id, which already goes through the shared boundary helper.Both fields now validate the same way their siblings do, and
verify_agent_idis canonicalised the same way on write (feat(goals): gate autonomous goal runs on a verifier and an evaluator #7785) (@DaBlitzStein)PUT /api/goals/{id}now actually stops its autonomous run, instead of being silently undone by the run itself.Deleting a goal has always stopped its run; updating one never did, and nothing else connected an operator's decision to the run's lifecycle — the runner only ever noticed by re-reading the goal document on its next tick.
That read cannot tell an operator apart from the
goal_updatetool the agent's own prompt tells it to call, so once a verifier was configured the runner correctly stopped treating a barestatus: completedas a reason to finish, and the operator's path went with it.An operator who pressed the dashboard's status button on a verified goal therefore watched it flip straight back to
in_progress, keep the incoherentprogress: 100that came with it, and spend the rest of its iteration budget on paid turns nobody had asked for.The two writers are now separated by which channel they use rather than by guessing from the stored value: an operator gets the run's real stop control, and an agent asserting completion in a document still has to get past the verifier.
The run also no longer writes its own status and progress over a goal that the same request has just written, so the choice an operator made mid-iteration survives the iteration already in flight.
A plain
POST /api/goals/{id}/stopis unaffected and still lands the interrupted iteration's progress, because it writes nothing to the goal there is anything to protect (feat(goals): gate autonomous goal runs on a verifier and an evaluator #7785) (@DaBlitzStein)loop_engineeringwith a verifier configured could still finish on work the verifier had just rejected, because completion was read offgoal.progress/goal.statusdirectly instead of off the verifier's own decision, and both fields have a second writer — thegoal_updatetool, which the agent's system prompt tells it to call independently of the text markers the verifier gate inspects.The runner now treats bare progress and a bare
status: completed, from either writer, as a completion signal only when no verifier is configured; a verified run requires the gate's ownGOAL_DONE-after-PASS branch to have actually run.Unrelated to that gate: an unreachable verifier — a deleted agent, a revoked key — fed no circuit breaker of its own, so a permanently dead one burned the whole iteration budget dispatching to it every round before reporting the cap as the reason it stopped, the exact waste the tick-failure breaker exists to prevent on the generator leg.
Five consecutive verifier failures now stop the run in
Stoppedwith the cause onlast_error(feat(goals): gate autonomous goal runs on a verifier and an evaluator #7785) (@DaBlitzStein)agent_spawn { profile }now checks the requested profile against the spawning agent's own[model.router_override].The same
allowed_profilesandcost_budgetthe per-turn router applies now also bind a named profile at spawn time, and a parent pinned withfixed = trueis refused outright — delegation was otherwise a way around every per-agent constraint the profile layer introduces, letting an agent budgeted atcheapspawn a helper on the most expensive profile in the catalog, billed to the same operator, with the parent's cap never applied.The refusal lists the profiles the spawning agent is permitted to use, so the calling agent can retry with a name that passes instead of guessing.
The override is also copied onto the agent that gets spawned, so the cap binds the whole delegation chain rather than only its first hop: a budgeted parent could otherwise spend its budget on a permitted child, hand that child
agent_spawn, and let it — born with no override, which every caller reads as unconstrained — spawn on the most expensive profile in the catalog for the same operator's bill.That also makes
fixed = truebind the subtree instead of merely refusing the pinned agent a profile of its own while it produces unpinned children that route freely.Failing to look the spawning agent's constraints up refuses the spawn only when a profile was actually requested, because a spawn that names no profile has no cap on this hop to enforce and would otherwise hard-fail against a caller-supplied agent id that never resolves, such as the REST tool endpoint's
agent_idor a deferred approval resumed after the requesting agent left the registry.The profile's model is now resolved through the live model-catalog alias table before being written to the child's manifest, matching what the per-turn router already does, because every builtin profile names an alias like
haikurather than a concrete model id and an unresolved alias would otherwise reach the provider verbatim and fail authentication on the agent's first turn.The same profile check, alias resolution and router-override guard now also apply to an ephemeral (
ephemeral: true) worker, which previously ignoredprofileentirely and ran on the default model with no indication the parameter had been dropped; an explicitmodeloverride on that path is refused when the parent isfixedor budgeted, since there is noprofilegate to check an arbitrary model id against.A profile whose provider has no configured credentials is refused before the agent is spawned, naming the missing environment variable, instead of being born and failing authentication on every turn thereafter.
A
profilevalue of the wrong JSON type — an object or array instead of a string — is now refused with an explanation rather than silently treated as absent.A profile refusal or an unknown-profile-name error is now reported as an invalid parameter or a permission denial instead of a generic upstream failure, so retry logic on the REST bridge does not treat an operator's own cost cap as a transient 5xx outage.
An agent whose router-override permits no catalog profile at all is now told so explicitly instead of getting an empty "Permitted profiles: ." list.
Pinning an ephemeral worker no longer leaves it running under the parent's context window, API key and base URL: those three describe the parent's own model, and an ephemeral worker inherits the parent's manifest wholesale, so a parent pinned to a million-token profile budgeted its Haiku worker at a million tokens — never compacting, and rejected by the provider — while a parent pinned to a custom OpenAI-compatible endpoint sent its worker's Anthropic requests to that endpoint with that key, moments after the spawn had verified that Anthropic's own credentials were present (feat(runtime): let agent_spawn pin the spawned agent to a model profile #7789) (@DaBlitzStein)
Clearing the parent's endpoint, key and context window off the primary model closed the obvious half of the inheritance, but
fallback_modelsis a sibling ofmodelon the manifest rather than a field inside it, and every entry in it carries its ownapi_key_envandbase_url.resolve_effective_fallbackstreats an agent's own list as the exclusive chain, so a parent whose chain falls back to its expensive model under a private proxy credential handed that entry straight to the worker: the first rate-limit on the cheap model promoted the worker onto the parent's model with the parent's key, past bothallowed_profilesand the cost budget the rest of the feature exists to enforce.It only fired on the second request of a run, which is why the first-request checks all looked correct.
max_tokensandextra_paramswent the same way for the same reason — the first reaches the wire with no clamp against the model's real ceiling, producing exactly the oversized-request rejection the context-window clear was added to prevent, and the second is flattened into the request body verbatim, so a parent on Qwen or OpenAI posted that provider's keys to Anthropic.max_output_tokensis deliberately left alone: it belongs to the same conceptual group but has no reader anywhere on the request path (feat(runtime): let agent_spawn pin the spawned agent to a model profile #7789) (@DaBlitzStein)load_pause_checkpointreportsNoneboth for "there is no checkpoint" and for "there is a row I could not parse" — a substrate read error is swallowed by its.ok().flatten(), and so is a row whoseagent_idis missing or malformed — and cancel used thatNoneto decide whether to delete anything.So a transient storage failure at cancel time left the checkpoint in place, and the next start silently resumed the run the operator had just cancelled, which is precisely the outcome the cancel path exists to prevent.
The delete is now unconditional and the read only decides what the call reports, which costs nothing because deleting an absent key was already a no-op. (feat(goals): add pause/resume and a configurable loop cadence #7973) (@DaBlitzStein)
POST /api/goalsnow treats a blankverify_agent_idas "not set" instead of 400ing as an invalid UUID, matching the existing blank-means-absent rule forparent_idandagent_id(Create a Goal, returns 404 #6562).A blank
evaluator_modelis now filtered the same wayparent_id/agent_idalready are on update, instead of being stored verbatim as an empty string that the field's own documentation says should read as "no evaluator configured". (feat(goals): add pause/resume and a configurable loop cadence #7973) (@DaBlitzStein)The pause/stop flags were only checked at the top of the run loop, so a
tick_interval_secsset close to its 24-hour maximum meant a requested pause could sit unobserved for up to a day; the inter-tick sleep now wakes every second to re-check them.The
GOAL_LEARNED:lessons a run collects before pausing are now carried into its resume checkpoint and threaded back into the resumed run's own accumulator, instead of resetting to nothing on every pause. (feat(goals): add pause/resume and a configurable loop cadence #7973) (@DaBlitzStein)POST /api/goals/{id}/resume(and/starton a paused goal, which auto-resumes the same way) now rejects an explicitmax_iterationsat or below the paused run's already-completed iteration count.Resuming with a cap that low used to immediately trip the iteration-cap check with no turn run and discard the checkpoint on the way out — including the learnings it carried — for a request that could never have advanced the run in the first place. (feat(goals): add pause/resume and a configurable loop cadence #7973) (@DaBlitzStein)
GoalRunner::startalready resolved the resumed iteration count into the run's observable state, but the run loop itself kept its own separate counter hardcoded to 0, so the loop's iteration cap and every progress write after the first tick counted from scratch — a run paused at iteration 30 under a cap of 100 got a fresh 100-iteration budget instead of the 70 remaining. (feat(goals): add pause/resume and a configurable loop cadence #7973) (@DaBlitzStein)verify_max_retriesnow survives a pause the same waymax_iterationsalready does.A run started with an explicit retry budget reported the compiled default instead once paused, and a bodyless
/resumere-budgeted it to that default rather than restoring the operator's own number — the checkpoint never carried the field. (feat(goals): add pause/resume and a configurable loop cadence #7973) (@DaBlitzStein)GOAL_LEARNED:lessons are no longer queued as a pending skill draft for an agent that never opted into the skill workshop.The workshop is default-off and opted into per agent (
agent.toml: [skill_workshop] enabled = true), but the goal runner's learnings hook queued a draft regardless, because nothing in that path read the setting.It now checks
enabledandauto_capture, the same gate every other automatic capture path in the workshop already applies. (feat(goals): add pause/resume and a configurable loop cadence #7973) (@DaBlitzStein)POST /api/goals/{id}/startand/resumenow validateverify_max_retriesthe same way they already validatemax_iterations: an out-of-range or wrongly-typed value gets a 400 naming the field, instead of a bareas u32cast that silently wrapped a value likeu32::MAX + 1down to something else entirely. (feat(goals): add pause/resume and a configurable loop cadence #7973) (@DaBlitzStein)The top-of-loop completion check read bare
goal.progress >= 100/status == Completedregardless of whether a verifier was configured, andgoal_update(a tool the agent's own system prompt tells it to call) writes those same fields directly, bypassing the marker parser the verifier gate actually inspects.The completion check now only accepts bare progress/status as done when no verifier is configured; a rejected iteration's progress is additionally clamped below the completion threshold as defense in depth.
Converges with the equivalent fix in PR feat(goals): gate autonomous goal runs on a verifier and an evaluator #7785, which found and closed the same second-writer bypass first. (feat(goals): add pause/resume and a configurable loop cadence #7973) (@DaBlitzStein)
librefang goal --watchnow recognizes apausedrun instead of treating it as unclassified state.The daemon can report a goal's run as
paused(an operator-triggered pause, resumable later), but the CLI's terminal-phase table had no entry for it, so--watchburned its bounded unobservable-poll retry budget and exited with a generic "gave up observing the run" message instead of reporting that the run was paused. (feat(goals): add pause/resume and a configurable loop cadence #7973) (@DaBlitzStein)Until now the operator's message was the last thing in the session: the turn returned an error to its caller, the history recorded nothing, and neither the open chat nor a reload explained where the answer went — the shape observed live when the provider circuit breaker opened mid-stream.
The session now keeps a short note saying the provider failed and that no response was produced.
The note is deliberately opaque and carries none of the driver's error text, which goes to the daemon log instead, because a provider error's
Displayroutinely drags along the endpoint URL, the model id and the upstream response body.The streaming path also persists the inbound message before it calls the provider at all, which is what the non-streaming path already did; a restart or a hang between the two used to lose the operator's message outright, and that is the path the dashboard takes. (feat(media): route inbound media by capability instead of dropping it #7989) (@DaBlitzStein)
POST /api/media/transcribeandPOST /api/agents/{id}/upload.Every SDK method was emitted through the same path, which serialises its argument with a JSON encoder and sends
Content-Type: application/json; both handlers read the body as bytes and reject anything that is not the content type they declare, so the shippedtranscribeAudio/uploadFilemethods returned 400 in every language, unconditionally.They now take the bytes and an optional content type, defaulting to the one the OpenAPI operation declares, and the generator derives that from the spec rather than from a list of special cases — an endpoint added later with a non-JSON body gets a working method without anyone remembering this. (feat(media): route inbound media by capability instead of dropping it #7989) (@DaBlitzStein)
POST /api/goals/{id}/resumereports404 Goal not foundfor a goal id that does not exist, exactly asPOST /api/goals/{id}/startalready did, instead of409 Conflict.The precondition that refuses a resume when there is no paused run ran before the goal was looked up, so an operator who mistyped an id was told the goal had nothing to resume and pointed at
/start— the same handler, which would itself have answered 404 (feat(dashboard): add pause/resume controls to the goals page #8029) (@DaBlitzStein)channel_sendon the kernel-internal channelswebui,cronandautonomousare now held in place by regression tests, after two branches rewrote the same block from different starting points and landed on opposite advice.The behaviour itself shipped with fix(runtime,kernel): error classification, system channel detection, and model sentinel tests #7995; what lands here are the seven tests that keep it from drifting back — that
webuiis never told its generated media reaches the browser on its own when the browser only ever sees what the reply text embeds, and that a background run is never told not to usechannel_sendat all when the only thing genuinely impossible there is replying into a sentinel channel that has no adapter.Each test names the sentence it pins, so a future edit to that wording fails with the reason rather than with a diff. (fix(runtime): stop suggesting channel_send to kernel-internal system channels #8149) (@DaBlitzStein)
librefang skill publishnow resolves its GitHub token the same way the HTTP routes do — the environment first, then the credential vault.It previously read
GITHUB_TOKEN/GH_TOKENfrom the environment only and exited with status 1, so the same daemon could promote a skill through the API and fail to publish one from the CLI on the same machine with the same token in the vault.The "no token" message now names both places a token can live instead of only the environment variables. (feat(skills): make the registry promotion GitHub target configurable, and align the CLI's token resolution #8179, Registry promotion: GitHub target, branch, commit author and API base URL are hardcoded or derived, not configurable #8163) (@DaBlitzStein)
The two surfaces each held their own
WORKFLOW_RUN_WAIT_MS— 90 s inlibrefang workflow run, 45 s in the Workflows screen — so a workflow that took 60 s completed from one and timed out from the other, with nothing on either screen to suggest the surface was the variable rather than the workflow.Each constant was correctly derived from its own caller's client timeout, which is why neither looked wrong in isolation: the TUI built that one request with a 60 s client, a local choice among the 5 s to 300 s timeouts it picks per call rather than a constraint.
Both the client timeout and the wait now come from one place, tied together by a compile-time assertion, and both surfaces build the request from the same helper so the query cannot disagree again. (fix(cli): give the TUI and the CLI the same workflow-run deadline #8317) (@houko)
Coverage was the only thing checked, so a key carried over with its English value passed exactly as well as a real translation — that is how eight Auxiliary-tab strings reached Ukrainian and Chinese in English and stayed green.
The failure is invisible in a way a missing key is not: a missing key renders as
[key]and looks broken, while English inside an otherwise translated screen reads as a deliberate choice, so nobody reports it.Values that are supposed to match English — brand names, shell commands a user copies verbatim, column headers that name an API field — are listed in a table with a written reason for each, and a second test fails when an entry there stops describing anything, so the exemption list cannot quietly become a rubber stamp. (test(cli): fail when a locale value is a copy of the English text #8313) (@houko)
[queue] max_depth_per_agent,max_depth_globalandtask_ttl_secsnow do what they have always been documented to do, instead of being three settings an operator could configure, read back from the API, and get nothing from.None of the three had an enforcement site anywhere in the codebase: every insert succeeded whatever the depth, no pending task ever expired, and because only terminal rows are pruned,
task_queuegrew for the life of the install — which then made every unpaged task-list request allocate one JSON object per row in the table.A post that would exceed a non-zero depth cap is refused with
429 Too Many Requestsand a message naming the cap it hit, counted in the same write transaction as the insert so two concurrent posts cannot both take the last slot.The per-agent cap is scoped to the assignee, and tasks with no assignee belong to the shared pool rather than to one bucket keyed on the empty string.
A task still unclaimed after
task_ttl_secsis moved tocancelledwith aresultnaming the setting, not deleted: a task an operator queued and has not yet staffed should not vanish without a record, andcancelledis the terminal status the dashboard, the status counts andtask_queue_retention_daysalready understand, so the row is reclaimed on the existing horizon and visible until then.A claimed task is in flight and is never expired this way.
Note that
task_ttl_secsships as3600, so an install that has never set it will start seeing hour-old unclaimed tasks cancelled; set it to0to keep the previous behaviour of never expiring anything.?limit=,?offset=and?assigned_to=onGET /api/tasksandGET /api/tasks/listare nowWHEREandLIMITclauses rather than aretainand atruncateover a fully materialised list, so asking for ten tasks costs ten rows instead of the whole table.totalkeeps its meaning — rows matching the filters, not the length of the page (fix(queue): enforce max_depth_per_agent, max_depth_global and task_ttl_secs #8373) (@houko)[docker] modeis set tonon_mainorall, and the documentation says the field is not implemented, instead of both presenting it as the switch that moves agents into containers.No execution path matches on the value:
shell_execandprocess_startrun as subprocesses on the daemon host whatever it says, and the only way into a container remains the separatedocker_exectool, which the model chooses for itself rather than the operator.What made this worth a warning rather than a doc note alone is that the rest of
[docker]is live —enabled,scope,reuse_cool_secs,idle_timeout_secs,max_age_secs, the image and the limits all govern the containersdocker_execcreates — so the section visibly works and the one field in it that does nothing looks like it works too.An operator who believed they had switched every agent onto OS-level isolation had switched nothing, with a clean boot and no log line anywhere.
DockerSandboxMode::is_wired_into_dispatchis the single source of truth, mirroring what tool_exec.kind and per-agent tool_exec_backend never reach tool dispatch, and the boot WARN the docs promise does not exist #8221 established for[tool_exec] kind— the adjacent knob with the same gap — and a test fails if a mode flips to wired without the warning and the docs being revisited (fix(docker): say out loud that [docker] mode routes nothing #8374) (@houko)Its tooltip used to report the run's status instead — "Running · iteration 3/10" — so the one control that stops an autonomous run described the run rather than the action, and the iteration counter it showed was already on the row beside it.
It now reads "Stop autonomous run" whether or not a run state has loaded, and the status-shaped
goals.run_activekey that fed it is gone from all five locales. (feat(tui): pause and resume a goal run from the Goals screen #8224) (@DaBlitzStein)Starting, stopping, pausing or resuming a run fires two requests on independent threads — one for the list, one for that goal's run state — and the list payload is built from stored goal documents, which never carry a phase.
Whichever landed second won, so about half the time the freshly fetched phase was overwritten with nothing, and
rdid it every time.The visible cost was on the pause key, which reads that phase to decide between pausing and resuming and does nothing at all when it is absent: a run could be paused and then not resumed, from the same screen that was still showing it as paused.
The list now merges by goal id and keeps a phase it already knows, rather than replacing the rows wholesale. (feat(tui): pause and resume a goal run from the Goals screen #8224) (@DaBlitzStein)
channel_sendto a channel namedcron,autonomousorwebuinow mirrors into the conversation the operator is looking at instead of the kernel's own internal session.Those three names are reserved because they are the kernel's system sessions, so every path that derives a channel-scoped session id renames an operator-supplied one to
ext-<name>first — except the mirror, which calledSessionId::for_sender_scopedirectly and so wrote the outbound message into the one session that must never carry channel traffic, while leaving it out of the chat that should show it.The guard moved from a
pub(super)helper on the kernel tolibrefang_channels::types::resolve_scope_channel, because being unreachable fromlibrefang-runtimeis what made it skippable in the first place. (fix(runtime): rename a reserved channel before deriving the channel_send mirror session #8316) (@houko)So does asking for one that exists under a different agent.
The kernel returned the miss as a string inside
LibreFangError::Internal, and the route helper typed only the two agent-shaped errors, so every other kernel error — including a plain bad id — became a server fault whose reason was then scrubbed out of the body.The scrub is right and stays, because the memory layer wraps every rusqlite error in that same variant and echoing one would leak SQL schema; what was wrong was calling a missing session an internal error in the first place.
A caller could not distinguish a typo from an outage, and a scripted client saw a retryable 5xx where the answer will never change.
The fix is typed rather than a match on the message text:
SessionNotFoundandResourceNotFoundalready existed and the sibling helper forKernelOpErroralready mapped both to 404, so this brings the outlier into line for the fifteen handlers that share it.That also fixes tool-level misses, which reach the same helper as
ResourceNotFoundand were 500 for the same reason.A session belonging to another agent is reported as not found rather than as a distinct wrong-owner error, so the answer does not confirm the session exists to someone who cannot read it. (fix(api): answer 404 for a missing session instead of 500 #8263) (@DaBlitzStein)
POST /api/mcp/servers/{name}/reconnectanswered 500 with a generic body for a server that never completed its handshake — sending the operator to look for a fault inside this daemon when what failed was an external dependency that did not answer, and giving them nothing to act on.The reason was in the daemon's journal the whole time, so learning it took an SSH session.
The answer is now 502 for a server that did not answer and 409 for one whose own stored configuration blocks the reconnect, and the body carries the failure class, the transport kind and the endpoint that was dialed with its arguments, path and query stripped — the parts that can hold a token.
Failed MCP connects are also recorded in the audit trail, which is what the dashboard's Logs page reads, so a server that will not start now leaves a trace on the screen an operator opens when something breaks instead of only in the system journal. (fix(api,kernel): answer 502 with a typed reason when an MCP server will not start #8271) (@DaBlitzStein)
text_to_speechnow takes its defaultoutput_formatfrom a new[tts] output_format, so a deployment whose channel accepts only Ogg/Opus voice notes sets it once instead of depending on the model to pass an optional argument on every call (on the paths that write to a workspace — a caller with no workspace root still gets the provider's bytes back unconverted).Only ElevenLabs had a configurable output format; every other provider was pinned to MP3, which a messaging channel rejects as a voice note — and because synthesis itself succeeded and wrote the file, the failure surfaced as a reply that never arrived, with nothing in the log naming the tool or the format.
The value is carried on the turn's
LoopOptionsrather than read off theTtsEnginehandle, because that handle is withheld whenever[tts] enabled = false— a state in which the tool still runs, on the media-driver path — and holds a boot-time clone besides; reading through it would have left the new key unreachable in the shipped default configuration and stale afterPOST /api/config/reload.build_reload_plannow classifies[tts]the way it classifiesregistry:enabledandoutput_formatare re-read per turn, everything else is captured inTtsEngineat boot and is reported restart-required instead of being answered with a false "effective on next message".The default is unchanged — unset still means
"mp3"— and an unrecognised value is reported byvalidate()at config load, since at the point of use it is indistinguishable from the default (fix(runtime): resolve the text_to_speech output format from config for every provider #8274) (@nevgenov)A half-open socket never fires
onclose, so the browser kept reportingreadyState === OPENand the frame was written into a connection whose bytes went nowhere: the turn spun forever, and reloading showed neither the question nor an answer, because nothing had reached the daemon to persist.The
visibilitychangewake-up added in fix(dashboard): recover ChatPage WS from retries-exhausted state on tab visible / online #4063 could not help, since it was gated on the retries-exhausted flag that only a disconnect the browser actually noticed can set — leaving the one case the listener existed for as the one case it could not act on.Coming back to the tab now probes the link with the
{"type":"ping"}/{"type":"pong"}exchange the daemon has answered since the socket was first written and no client had ever sent, and hands a socket that does not answer to the reconnect path that already exists.A probe is skipped while a turn is in flight, because the daemon stops reading the socket for the duration of a turn and could not answer one. (fix(dashboard): probe the chat socket on tab return instead of trusting readyState #8275) (@DaBlitzStein)
None of the three endpoints ever sent a Ping, so the only thing that could discover a dead peer was a failing write — and a connection sitting idle between turns, which is where a chat socket spends most of its life, is never written to at all.
Ten daemon lifetimes on a production host recorded 55
client_closedisconnects, onesend_error, onereceive_errorand not a singleidle_timeout: the one detection that existed worked, and only ever fired when there was outbound traffic.A half-open socket left by a suspend, a wifi roam or an expired NAT mapping was therefore held until the idle timeout, which deployments routinely set to hours.
The terminal socket was the worst affected, because its idle timer is reset by PTY output as well as by client input — so a shell that keeps printing kept a dead peer's child process, tmux window and connection slot alive indefinitely.
Detection costs at most two intervals and is tuned with
rate_limit.ws_ping_interval_secs(default 30 s,0disables); an answered Ping deliberately does not count as activity, sows_idle_timeout_secsstill fires on a genuinely idle browser tab. (fix(api): ping a silent WebSocket peer so a dead one is detected #8278) (@DaBlitzStein)[tts] provider, the[tts.google]block and[tts.elevenlabs] output_formatnow apply on the media-driver path, so they reach a deployment running the shipped[tts] enabled = falsedefault instead of being silently inert on it.text_to_speechis registered unconditionally and reaches providers throughMediaDriverCache, but those three reads came off theTtsEnginehandle, which the agent loop lends only whenenabled = true— so on the default configuration the tool worked while most of its own configuration section did nothing.An operator who set
[tts.google] language_codeand leftenabled = false, reasonably reading it as "enable the TTS feature" since the tool already worked for them, goten-USwith nothing in the log; one who named a provider was auto-detected onto a different one and billed there.A pinned provider that turns out not to be configured for text-to-speech now degrades to capability detection with a warning, rather than failing the call:
get_or_createdoes not screen on credentials the waydetect_for_capabilitydoes, so honouring the pin without a fallback would have converted a working auto-detection into a hard failure for precisely the deployments this fixes.The provider-specific overrides also key off the driver that will actually serve the request rather than the name that was asked for, which is what makes them apply when detection picks Google rather than only when Google is named.
The Google voice override stays unconditional rather than gaining the
is_none()guard its ElevenLabs neighbour has: for Google an OpenAI-style voice name such asalloyis not a preference to respect but a request that fails, and adding the guard would have taken that safety net away from every deployment that has it today.All three keys move to the reload-classification's live half as a result, and
MediaDriverCachegains a test-only seeding point so the engine-less media-driver shape — which had no coverage at all, and is where this and text_to_speech: output_format hardcoded to mp3 for every provider except ElevenLabs — voice notes fail silently #8272 both live — is now asserted against a stub driver that records what the tool asked the provider for (fix(tts): apply [tts] on the media-driver path, not only through the engine #8375) (@houko)librefang doctoroutput is readable again: 71zh-CNvalues were generated from their own key names rather than translated, and have been rewritten.CLI is up to daterendered ascliuptodate,Database status: { $status }asdb状态fail失败:{ $status }, and theChannel Integrations:section heading asdoctorsection频道.Two of them changed behaviour rather than only readability — the
[Y/n]was missing from bothdoctorconfirmation prompts, so a Chinese user was asked a yes/no question with no indication of what to type or which answer was the default, and the.env file not foundwarning dropped thelibrefang config set-keycommand that resolves it.Korean and Ukrainian were unaffected; the values trace to a single bulk import in feat(cli): localize TUI Onboarding Wizard and Agents screen #6253. (fix(cli): translate the 71 zh-CN values that were generated from their key names #8315) (@houko)
The rule that a binary will not touch a database newer than itself is the right one — it cannot know what a later step did, so writing over it would corrupt whatever that step added — but the effect was that a machine which had run an unreleased build was locked out of every release binary, with a boot loop and "Downgrade is not supported" as its only signal.
The two steps involved add nothing new: one is a column a later-numbered migration already creates, and the other is a table created only if it is missing, so a database that never saw those builds is unchanged.
Changelog truncated — GitHub caps a PR body at 65,536 characters. The full section is in CHANGELOG.md.
Full diff: v2026.9.14...v2026.9.19