Repository navigation
feat(config,drivers): per-agent and per-task reasoning mode for DeepSeek V4 and OpenRouter - #8083
Conversation
Add `reasoning_mode = "none" | "low" | "high" | "max"` as a first-class
setting on the existing `[thinking]` table, so it is settable globally in
`config.toml`, per agent in `agent.toml`, and per task on both message
endpoints. Resolution order is per-call > per-agent > global > compiled
default, matching the shape already documented for `max_history_messages`
and `session_mode`.
Before this, the choice could not be expressed at all, and for DeepSeek
V4 — direct or via OpenRouter — the driver sent no reasoning fields
whatsoever, so every request ran at the provider default of thinking-on
at `high`. Prompt-level workarounds are not a substitute: the reasoning
tokens are still generated, still consumed from the output budget, and
still billed.
The mode is translated per provider in one resolver,
`reasoning_wire_fields`, keyed off `base_url` rather than the model id:
* DeepSeek V4 direct — `thinking: {"type": "disabled"}` for `none`,
otherwise `thinking: {"type": "enabled"}` plus `reasoning_effort`
(`low` / `high` / `max`).
* OpenRouter-routed — the nested `reasoning: {effort}` form, with
`max` as `xhigh`. Never together with top-level `reasoning_effort`,
which OpenRouter answers with HTTP 400.
* Other OpenAI-compatible endpoints — top-level `reasoning_effort`,
with `none` expressed by omission and `max` clamped to `high`.
`budget_tokens` remains the fallback bucket, and an unset
`reasoning_mode` leaves every wire shape byte-for-byte unchanged.
Also fixes the SSE streaming endpoint, which passed a hardcoded `None`
into the kernel and so dropped the per-call thinking override entirely,
and extends the #7769 strip-and-retry to clear both spellings of the
reasoning control.
The per-call key is a new `reasoning_mode` field rather than a widened
`thinking` boolean: the boolean is in the wire contract of every existing
client, and `thinking: false` only omits the opt-in while
`reasoning_mode: "none"` sends the provider's explicit non-think toggle.
Dashboard and TUI exposure is deliberately left out.
Refs #7946
An extra_body.reasoning_effort override on an OpenRouter-routed request left the nested `reasoning` object reasoning_wire_fields had already set, so the merged body could carry both spellings — the exact combination OpenRouter answers HTTP 400 to. merge_extra_body now clears whichever spelling an override key replaces. Also fills in the changelog fragment's literal (#PR) placeholder with the actual PR number.
houko
left a comment
There was a problem hiding this comment.
Automated review pass against CLAUDE.md's rules.
Pushed two mechanical fixes directly to this branch (commit 78e19a7):
merge_extra_bodydidn't clear the counterpart reasoning key on anextra_bodyoverride, so an OpenRouter-routed request withextra_params.reasoning_effortset could end up carrying bothreasoningandreasoning_efforton the wire — the exact combination this PR's own invariant says OpenRouter 400s on. Added a regression test (extra_body_reasoning_effort_override_clears_the_nested_openrouter_spelling) alongside the existing coverage.- The changelog fragment ended with a literal, unfilled
(#PR)placeholder instead of(#8083)— perchangelog.d/README.mdthis means the PR would keep its generated release-notes line and appear twice.
Two remaining items left as inline comments below since they need judgment calls I can't make confidently from here (real provider error text, and a design decision on shared helpers):
- The strip-and-retry rejection matcher may not recognize an OpenRouter gateway's rejection of the new nested
reasoningfield, which would silently break the #7769 self-healing retry for that dialect. - A minor duplicated "is this OpenRouter" substring check between
reasoning_wire_family()andshared_guard_provider().
cargo check -p librefang-llm-drivers --lib, cargo clippy -p librefang-llm-drivers --all-targets -- -D warnings, and cargo test -p librefang-llm-drivers --lib all pass with the fix.
Generated by Claude Code
| if status == 400 | ||
| && oai_request.reasoning_effort.is_some() | ||
| && (oai_request.reasoning_effort.is_some() || oai_request.reasoning.is_some()) | ||
| && crate::llm_driver::llm_errors::is_unsupported_reasoning_effort_error(&body) |
There was a problem hiding this comment.
The comment above says "Both spellings are stripped (#7946)", but the gate itself (is_unsupported_reasoning_effort_error, in crates/librefang-llm-driver/src/llm_errors.rs, unmodified by this PR) still requires the literal substring "reasoning_effort" in the rejection body. For an OpenRouter-routed request that now sends the nested reasoning: {"effort": ...} object instead, a gateway rejecting that field is more likely to name it "reasoning" than "reasoning_effort" in its error text — in which case this check returns false, the field is never cleared, record_reasoning_effort_rejected (line below) never fires, and the #7769 negative cache never populates for that model. The agent would then fail identically on every subsequent turn instead of self-healing after one retry, unlike the sibling temperature/max_tokens strips just above.
Same gate is duplicated at line ~2112 for the streaming path.
Worth confirming what real OpenRouter-adjacent gateways actually say in a 400 body for the nested field, and widening is_unsupported_reasoning_effort_error (or adding a sibling matcher) to catch the bare "reasoning" rejection if so. Leaving this as a comment rather than a same-PR fix since the correct pattern depends on real provider error text.
Generated by Claude Code
| /// `reasoning: {"effort": …}` form or top-level `reasoning_effort` is | ||
| /// accepted. `openrouter/deepseek/deepseek-v4-pro` and a direct | ||
| /// `deepseek-v4-pro` are the same model behind two different dialects. | ||
| fn reasoning_wire_family(&self) -> ReasoningWireFamily { |
There was a problem hiding this comment.
Minor: reasoning_wire_family() re-implements shared_guard_provider()'s base_url.to_ascii_lowercase().contains("openrouter") check a few lines below, for a different purpose (wire dialect vs. rate-limit bucket tag). Not a bug today, but a future host alias added to one and not the other would silently desync which dialect a request gets from which rate-limit bucket it's grouped under. Might be worth factoring the "is this OpenRouter" test into one shared helper both call — not blocking.
Generated by Claude Code
Deploying librefang-docs with
|
| Latest commit: |
4943d18
|
| Status: | ✅ Deploy successful! |
| Preview URL: | https://8616e155.librefang-docs.pages.dev |
| Branch Preview URL: | https://feat-7946-per-agent-reasonin.librefang-docs.pages.dev |
`GET /api/config` omitted `thinking.reasoning_mode` while `POST /api/config/set` accepted it, which the `every_writable_config_leaf_is_readable` parity guard (#6596) turns into a test failure — and which would have made the dashboard render the field blank and read a successful save back as "not configured". `redacted_config_json` now emits it, through `ReasoningMode`'s serde form rather than `Debug` so the value matches the options the dashboard's dropdown offers. DeepSeek R1 routed through OpenRouter would have started receiving a budget-derived `reasoning.effort` it never received before, because the OpenRouter family branch ran ahead of the `Strip` echo-policy check. R1 reasons unconditionally and exposes no wire toggle, so the field changes the payload without changing the model. The `Strip` check now short-circuits first, in both dialects. Also regenerates the `KernelConfig` JSON Schema golden fixture for the new `ReasoningMode` definition and the `ThinkingConfig` property. Refs #7946
|
Two follow-up commits are now on the branch (tip
|
…#8083 `reasoning_mode = "none"` turned reasoning ON for the budget-based drivers. `apply_thinking_override` has to materialise a `ThinkingConfig` to carry the mode, and its default `budget_tokens` is 10_000. Anthropic read that as "enable extended thinking" and inflated `max_tokens` by `budget + 1024`; Ollama's `think: Some(request.thinking.is_some())` read it as `think: true`. An agent that asked not to reason reasoned, and was billed for it. Both drivers now treat `reasoning_mode = "none"` as off. The strip-and-retry never fired for the nested OpenRouter spelling. `is_unsupported_reasoning_effort_error` still required the literal `reasoning_effort` in the body, so a gateway answering `unsupported parameter: reasoning` matched nothing: no strip, no retry, no negative-cache entry, and the same 400 on every subsequent turn. Both names are recognised now; `reasoning_content` stays excluded, because DeepSeek's echo-it-back demand is a content requirement that stripping a control would not answer. `thinking: true` was a silent no-op against an inherited `reasoning_mode = "none"`. The boolean documents itself as "force thinking on even if the manifest has it off", and a non-think mode is exactly that off-state, so `Enable` now clears it. A graded mode is left alone. An `extra_body` supplying both spellings ended up with neither. Each guard removed the counterpart because the other key was present, so an operator who overrode both lost both and the model ran at its provider default with nothing in the payload to explain why. Left as-is deliberately: the `effort_rejected` short-circuit in `reasoning_wire_fields` returns ahead of the `Echo` branch, so a cached rejection also suppresses V4's `thinking: {"type": "disabled"}` gate. Emitting it anyway risks a permanent 400 on a gateway that does not understand `thinking` either — the retry loop clears only the two effort fields — so the conservative direction was kept and documented. Verification: cargo check --workspace --lib; cargo clippy -p librefang-llm-driver -p librefang-llm-drivers --all-targets -- -D warnings; cargo clippy -p librefang-kernel --lib -- -D warnings; cargo test -p librefang-llm-driver --lib (68 passed); cargo test -p librefang-llm-drivers --lib (670 passed); cargo test -p librefang-kernel --lib (1720 passed).
The visual agent editor re-emits `[thinking]` from the two fields it has widgets for, and #7946 adds a third. `parseManifestToml` read `budget_tokens` and `stream_thinking` into the form and dropped everything else on the floor, so opening any agent in the editor and pressing save rewrote its `agent.toml` without `reasoning_mode` — the per-agent knob the feature exists to provide, destroyed by the surface most operators will use to set it. `[capabilities]`, `[model]` and `[resources]` already guard this with a `stripKnown` / extras round-trip; `[thinking]` had no such slot because until now the form was a complete description of the table. It gets one, populated on parse and re-emitted inside the section on serialize, and dropped along with the rest when the user unticks the section. `generateManifestMarkdown` carried an inline `ManifestExtras` literal that silently drifts every time the interface grows; it now calls `emptyManifestExtras()` so the compiler catches the next one, which is how this was caught. The four configuration pages still said the Anthropic / Gemini / Ollama drivers "continue to use `budget_tokens`". That stopped being true when those drivers learned to honour `reasoning_mode = "none"`, and a reader trusting the old sentence would conclude the non-think rung does nothing on Claude. Corrected in both the English pages and their `zh` mirrors, matching the wording already in `docs/architecture/reasoning-mode-resolution.md`. No changelog fragment: the loss is only reachable through #7946's own field, which has not shipped, so there is nothing for a release reader to have been bitten by. Verification: `pnpm typecheck` clean, `pnpm eslint` clean on the changed files, `pnpm vitest run` 1674 passed across 200 files. `round_trips an unknown [thinking] key such as reasoning_mode` was confirmed to fail against the unmodified source before the fix was applied. Refs #7946
Review of #8083 — two follow-up commits pushed (
|
Closes #7946.
Design decisions
Config surface.
reasoning_mode = "none" | "low" | "high" | "max"was added as a field on the existingThinkingConfig, not as a new top-level knob. That one struct is already the global[thinking]table inconfig.tomland the per-agent[thinking]table inagent.toml, so a single field gets both rungs and needs no new plumbing between the kernel and the driver. Per CLAUDE.md #5476 the per-agent knob is inagent.toml—KernelConfighas noagentsfield, so[agents.<name>.thinking]inconfig.tomlwould parse and reach no manifest.Per-call field shape: a new
reasoning_modekey, not a widenedthinking. The boolean is in the wire contract of every existing client and a JSONtruecannot grow a third value without breaking them. The two are also not redundant:thinking: falsemerely omits the reasoning opt-in, whilereasoning_mode: "none"sends the provider's explicit non-think toggle — the difference between "we did not ask it to think" and "we asked it not to", which is the entire point on a model that reasons by default. When both keys are presentreasoning_modewins. Below the HTTP boundary the two collapse into oneThinkingOverridevalue so nothing downstream threads two parameters that can disagree.Resolution order: per-call > per-agent > global > compiled default, the same shape CLAUDE.md documents for
max_history_messagesandsession_mode. Mechanically: the manifest carries the agent's own table; the kernel backfills the global table only when the manifest has none (so a per-agent table wins by suppressing the backfill, not by merging);apply_thinking_overridestamps the per-call value on last; whatever survives is copied onto theCompletionRequest.mediumis not a settable mode. It remains reachable only as thebudget_tokensfallback bucket. A fifth rung whose only distinct meaning is "DeepSeek folds it intohigh" is a knob that does nothing on the provider that motivated the feature. TOML parsing rejects it explicitly rather than silently falling back.Unset
reasoning_modechanges nothing. A DeepSeek V4 body with no mode is byte-for-byte what it was before this PR, and thebudget_tokensbucket remains the fallback everywhere it applied. DeepSeek ignoresbudget_tokens, so deriving an effort from the budget for V4 would have changed the wire shape of every existing V4 request without changing what the model does. Opting in is how behaviour changes.Deliberately left to the UI: the dashboard agent-edit form and the TUI do not surface
reasoning_mode. The issue lists UI exposure as optional and it is omitted to keep this PR reviewable.Per-provider wire mapping
Resolved in one function,
reasoning_wire_fieldsinopenai.rs. The family comes frombase_url, not the model id —openrouter/deepseek/deepseek-v4-proand a directdeepseek-v4-proare the same model behind two different dialects.Echopolicy)Nonepolicy)nonethinking: {"type":"disabled"}reasoning: {"effort":"none"}lowthinking: {"type":"enabled"}+reasoning_effort: "low"reasoning: {"effort":"low"}reasoning_effort: "low"highthinking: {"type":"enabled"}+reasoning_effort: "high"reasoning: {"effort":"high"}reasoning_effort: "high"maxthinking: {"type":"enabled"}+reasoning_effort: "max"reasoning: {"effort":"xhigh"}reasoning_effort: "high"reasoning: {"effort": <budget bucket>}reasoning_effort: <budget bucket>reasoning_effortandreasoning.effort— OpenRouter answers HTTP 400 to one that does. The invariant is asserted exhaustively over every (family × echo policy × mode × budget × rejected) combination, not case by case, because a future branch that sets the wrong pair would otherwise surface only as a production 400 on one provider.noneis expressed by omission on generic endpoints because OpenAI answers 400 to the literalreasoning_effort: "none";maxis clamped tohighthere because the OpenAI schema has no rung above it.openai.rsclaiming "this API family rejects requests with unexpected reasoning fields" predated V4 and is corrected.EmptyStringpolicy) keeps forcingthinking: {"type":"disabled"}regardless of mode — that is a multi-turntool_callscorrectness requirement, not a preference.Strip) still gets nothing, in either dialect. TheStripcheck short-circuits ahead of the OpenRouter branch on purpose: without it, an R1 route through OpenRouter with a configuredbudget_tokenswould have started receiving areasoning.effortit never got before, changing the payload without changing the model (R1 reasons unconditionally and has no toggle). Caught on self-review and pinned bydeepseek_r1_gets_no_reasoning_control_in_either_dialect.Changes
librefang-types: newReasoningModeenum (serdesnake_case,Serialize/Deserialize/JsonSchema,as_str/Display); newreasoning_mode: Option<ReasoningMode>field onThinkingConfigwith#[serde(default)]and aDefaultimpl entry; newThinkingOverride(Inherit/Enable/Disable/Mode) withresolve(Option<bool>, Option<ReasoningMode>)andFrom<Option<bool>>.librefang-llm-drivers:reasoning_wire_fieldsplus the three per-family vocabulary functions; newreasoningfield onOaiRequestfor OpenRouter's nested object;reasoning_wire_family()derived frombase_url; the three reasoning fields of the outbound body now all come from the one resolver.librefang-llm-drivers: the bug(llm): a gateway that rejectsreasoning_effortfails every turn and then opens the provider circuit breaker #7769 strip-and-retry in bothcomplete()andstream()now clears both spellings, and the negative cache suppresses both, since an OpenRouter-routed body carries the nested object rather than the top-level field.librefang-kernel:thinking_overridethreaded asThinkingOverrideinstead ofOption<bool>through the message and streaming dispatch chains; pre-existing public entry points that take the boolean (the/thinkchannel toggle, WebSocket) keep their signature and convert at the boundary viaFrom.apply_thinking_overridehandles the four cases.librefang-kernel:send_message_streaming_with_incognitogained the override parameter. It previously passed a hardcodedNone, so the SSE streaming endpoint silently dropped the per-callthinkingoverride entirely — a per-task mode wired only into the buffered route would have been invisible to any streaming UI. Both routes now honour both keys.librefang-api:reasoning_modeonMessageRequest; bothsend_messageandsend_message_streambuild the override withThinkingOverride::resolve.librefang-api:redacted_config_jsonnow emitsthinking.reasoning_mode, soGET /api/configexposes whatPOST /api/config/setaccepts. Theevery_writable_config_leaf_is_readableparity guard (GET /api/configomits config fields thatPUTaccepts — write-only settings are invisible in the dashboard #6596) caught the omission — without it the dashboard would render the field blank and read a successful save back as "not configured".crates/librefang-api/tests/fixtures/kernel_config_schema.golden.json— the newReasoningModedefinition plus theThinkingConfigproperty.docs/architecture/reasoning-mode-resolution.md;[thinking]sections updated indocs/src/app/configuration/page.mdx,docs/src/app/configuration/features/page.mdxand bothzhmirrors;librefang.toml.example(with itsxtask/baselines/config.sha256recomputed).changelog.d/added/7946-per-agent-reasoning-mode.md.Verification
Scoped to the four crates this PR touches.
cargo clippy --workspace --all-targetswas NOT run: the shared host ran out of disk during it (No space left on devicewhile writing librefang-api's query cache), and the coordinator asked that it not be retried. CI runs the workspace-wide lane.New tests:
Wire shape —
crates/librefang-llm-drivers/src/drivers/openai.rs, each asserting the serialized body and what is absent as well as what is present:deepseek_v4_direct_none_mode_disables_thinking_on_the_wiredeepseek_v4_direct_graded_modes_enable_thinking_with_effort(low / high / max)deepseek_v4_without_a_mode_keeps_the_pre_7946_wire_shapeopenrouter_routed_deepseek_uses_the_nested_reasoning_object(all four modes,max→xhigh)openrouter_falls_back_to_the_budget_bucket_without_a_modegeneric_openai_compatible_model_maps_modes_to_top_level_effortexplicit_mode_wins_over_the_budget_bucketkimi_disable_wins_over_an_explicit_reasoning_modea_rejected_reasoning_control_suppresses_both_spellingsdeepseek_r1_gets_no_reasoning_control_in_either_dialectno_payload_ever_carries_both_reasoning_controls— the exhaustive mutual-exclusion guardResolution order —
crates/librefang-kernel/src/kernel/tests.rs:test_reasoning_mode_global_reaches_an_agent_that_declares_nothingtest_reasoning_mode_per_agent_beats_globaltest_reasoning_mode_per_call_beats_per_agent_beats_globaltest_reasoning_mode_per_call_creates_thinking_config_when_absenttest_disable_and_mode_none_are_not_the_same_thingConfig / manifest surface —
librefang-types:test_thinking_config_parses_reasoning_mode_from_toml,test_thinking_config_reasoning_mode_defaults_to_none,test_thinking_config_rejects_an_unknown_reasoning_mode,test_reasoning_mode_serde_spelling_matches_as_str,test_thinking_override_resolve_precedencetest_manifest_thinking_reasoning_mode_round_trips_through_toml,test_manifest_thinking_without_reasoning_mode_stays_absentAPI —
crates/librefang-api/tests/agents_routes_integration.rs,#[tokio::test]against the production router:test_reasoning_mode_field_accepted_by_message_endpointtest_invalid_reasoning_mode_is_rejected— the conclusive one.MessageRequestdoes not deny unknown fields, so "accepted" alone would also pass for a typo'd key; a"banana"mode can only 422 if serde is genuinely parsing the field intoReasoningMode.test_reasoning_mode_omitted_and_legacy_thinking_boolean_still_worktest_reasoning_mode_accepted_by_streaming_message_endpointtest_per_call_reasoning_mode_wins_over_the_legacy_boolean— asserts the override the handler builds, i.e. the injection site, per CLAUDE.md's warning that anOption::Nonedefault compiles silently while disabling the featurePrompt-ordering determinism (#3298) does not apply: nothing added here reaches a prompt. The one map-shaped surface in the outbound body,
extra_body, was already aBTreeMap, and the newreasoningfield is a fixed-shape single-key object.Human-only verification (live daemon + real provider)
Not run here — live daemon plus a real LLM is human-only per CLAUDE.md. The wire shape is pinned by unit tests, but only a real endpoint proves the provider accepts the payload.
Start the daemon the usual way (
librefang+ thestartsubcommand) with a real key configured, then:DeepSeek V4 direct — non-think, then
max:To confirm the bytes rather than the outcome, run the daemon with
RUST_LOG=librefang_llm_drivers=traceand grep the request bodies forthinking/reasoning:none→"thinking":{"type":"disabled"}and noreasoning_effortmax→"thinking":{"type":"enabled"}and"reasoning_effort":"max"OpenRouter-routed — the mutual-exclusion case, which is the one that 400s if it regresses. Point an agent at
openrouter/deepseek/deepseek-v4-prowith a realOPENROUTER_API_KEY:Per-agent and global rungs — put
[thinking] reasoning_mode = "none"in the agent'sagent.tomland[thinking] reasoning_mode = "max"inconfig.toml, then:Streaming route — the same body against
/api/agents/{id}/message/stream, to confirm the per-call override now reaches the SSE path (it did not before this PR).Deferred / out of scope
budget_tokensand ignorereasoning_mode. Their wire control is a token budget rather than an effort enum, so honouring the mode there means inventing a mode→budget mapping that would silently change the budget existing Claude users already configured and pay for. That is a different provider family and a different behavioural risk than the issue describes, and it wants its own decision rather than a mapping picked here — flagging rather than deciding unilaterally, per CLAUDE.md. Stated explicitly in the architecture page's "Providers that ignore the mode" table so it is a documented boundary, not a surprise./thinkboolean only. Adding a mode there means changing the WS message envelope, which is the UI work above.openapi.jsonand bothxtask/baselines/*.sha256files are regenerated and committed here (cargo xtask codegen --openapi's underlying test, thenshasum -a 256— the same bytesschema-check genwrites).scripts/codegen-sdks.pyproduced no SDK diff. CI'sopenapi-driftlane should therefore find everything in sync.ConfigPage.tsxrenders schema-driven config fields throught(config.fld_${key}, fieldLabelFallback(key)), so[thinking] reasoning_modedegrades to a humanized "Reasoning Mode" label rather than breaking; adding a key toen.jsonalone would triplocale-parity.test.ts.dashboard/src/lib/agentManifest.tsliststhinkinginFORM_TOP_LEVEL_KEYS, soserializeManifestFormrebuilds the entire[thinking]table from its form model (agentManifest.ts:78-82type,:537-542write,:999-1003parse). A key the form model does not hold is a key a dashboard agent-edit silently deletes fromagent.toml— so an agent pinned toreasoning_mode = "none"would quietly revert to the provider default of thinking-on, and be billed for the reasoning. The fix is ~6 lines of preservation (carry it as an opaque string through parse and write, no UI), not the optional form field. I wrote it, then dropped it on instruction to keep this PR to the config-plus-wire surface, and becausedashboard/node_modulesis absent on this host so the vitest lane could not be run locally. Worth doing in the same cycle as the UI work, ahead of anyone setting the key and then editing that agent in the dashboard.