Skip to content

feat(config,drivers): per-agent and per-task reasoning mode for DeepSeek V4 and OpenRouter - #8083

Merged
houko merged 7 commits into
mainfrom
feat/7946-per-agent-reasoning-mode
Sep 1, 2026
Merged

houko merged 7 commits into
mainfrom
feat/7946-per-agent-reasoning-mode

Conversation

@houko

@houko houko commented Aug 31, 2026 •

Copy link
Copy Markdown
Contributor

Closes #7946.

Design decisions

Config surface. reasoning_mode = "none" | "low" | "high" | "max" was added as a field on the existing ThinkingConfig, not as a new top-level knob. That one struct is already the global [thinking] table in config.toml and the per-agent [thinking] table in agent.toml, so a single field gets both rungs and needs no new plumbing between the kernel and the driver. Per CLAUDE.md #5476 the per-agent knob is in agent.toml — KernelConfig has no agents field, so [agents.<name>.thinking] in config.toml would parse and reach no manifest.

Per-call field shape: a new reasoning_mode key, not a widened thinking. The boolean is in the wire contract of every existing client and a JSON true cannot grow a third value without breaking them. The two are also not redundant: thinking: false merely omits the reasoning opt-in, while reasoning_mode: "none" sends the provider's explicit non-think toggle — the difference between "we did not ask it to think" and "we asked it not to", which is the entire point on a model that reasons by default. When both keys are present reasoning_mode wins. Below the HTTP boundary the two collapse into one ThinkingOverride value so nothing downstream threads two parameters that can disagree.

Resolution order: per-call > per-agent > global > compiled default, the same shape CLAUDE.md documents for max_history_messages and session_mode. Mechanically: the manifest carries the agent's own table; the kernel backfills the global table only when the manifest has none (so a per-agent table wins by suppressing the backfill, not by merging); apply_thinking_override stamps the per-call value on last; whatever survives is copied onto the CompletionRequest.

medium is not a settable mode. It remains reachable only as the budget_tokens fallback bucket. A fifth rung whose only distinct meaning is "DeepSeek folds it into high" is a knob that does nothing on the provider that motivated the feature. TOML parsing rejects it explicitly rather than silently falling back.

Unset reasoning_mode changes nothing. A DeepSeek V4 body with no mode is byte-for-byte what it was before this PR, and the budget_tokens bucket remains the fallback everywhere it applied. DeepSeek ignores budget_tokens, so deriving an effort from the budget for V4 would have changed the wire shape of every existing V4 request without changing what the model does. Opting in is how behaviour changes.

Deliberately left to the UI: the dashboard agent-edit form and the TUI do not surface reasoning_mode. The issue lists UI exposure as optional and it is omitted to keep this PR reviewable.

Per-provider wire mapping

Resolved in one function, reasoning_wire_fields in openai.rs. The family comes from base_url, not the model id — openrouter/deepseek/deepseek-v4-pro and a direct deepseek-v4-pro are the same model behind two different dialects.

Mode DeepSeek V4 direct (Echo policy) OpenRouter-routed Generic OpenAI-compatible (None policy)
none thinking: {"type":"disabled"} reasoning: {"effort":"none"} field omitted
low thinking: {"type":"enabled"} + reasoning_effort: "low" reasoning: {"effort":"low"} reasoning_effort: "low"
high thinking: {"type":"enabled"} + reasoning_effort: "high" reasoning: {"effort":"high"} reasoning_effort: "high"
max thinking: {"type":"enabled"} + reasoning_effort: "max" reasoning: {"effort":"xhigh"} reasoning_effort: "high"
unset nothing (unchanged) reasoning: {"effort": <budget bucket>} reasoning_effort: <budget bucket>
  • No payload ever carries both reasoning_effort and reasoning.effort — OpenRouter answers HTTP 400 to one that does. The invariant is asserted exhaustively over every (family × echo policy × mode × budget × rejected) combination, not case by case, because a future branch that sets the wrong pair would otherwise surface only as a production 400 on one provider.
  • none is expressed by omission on generic endpoints because OpenAI answers 400 to the literal reasoning_effort: "none"; max is clamped to high there because the OpenAI schema has no rung above it.
  • The stale comment in openai.rs claiming "this API family rejects requests with unexpected reasoning fields" predated V4 and is corrected.
  • Kimi/Moonshot (EmptyString policy) keeps forcing thinking: {"type":"disabled"} regardless of mode — that is a multi-turn tool_calls correctness requirement, not a preference.
  • DeepSeek R1 (Strip) still gets nothing, in either dialect. The Strip check short-circuits ahead of the OpenRouter branch on purpose: without it, an R1 route through OpenRouter with a configured budget_tokens would have started receiving a reasoning.effort it never got before, changing the payload without changing the model (R1 reasons unconditionally and has no toggle). Caught on self-review and pinned by deepseek_r1_gets_no_reasoning_control_in_either_dialect.

Changes

  • librefang-types: new ReasoningMode enum (serde snake_case, Serialize/Deserialize/JsonSchema, as_str/Display); new reasoning_mode: Option<ReasoningMode> field on ThinkingConfig with #[serde(default)] and a Default impl entry; new ThinkingOverride (Inherit / Enable / Disable / Mode) with resolve(Option<bool>, Option<ReasoningMode>) and From<Option<bool>>.
  • librefang-llm-drivers: reasoning_wire_fields plus the three per-family vocabulary functions; new reasoning field on OaiRequest for OpenRouter's nested object; reasoning_wire_family() derived from base_url; the three reasoning fields of the outbound body now all come from the one resolver.
  • librefang-llm-drivers: the bug(llm): a gateway that rejects reasoning_effort fails every turn and then opens the provider circuit breaker #7769 strip-and-retry in both complete() and stream() now clears both spellings, and the negative cache suppresses both, since an OpenRouter-routed body carries the nested object rather than the top-level field.
  • librefang-kernel: thinking_override threaded as ThinkingOverride instead of Option<bool> through the message and streaming dispatch chains; pre-existing public entry points that take the boolean (the /think channel toggle, WebSocket) keep their signature and convert at the boundary via From. apply_thinking_override handles the four cases.
  • librefang-kernel: send_message_streaming_with_incognito gained the override parameter. It previously passed a hardcoded None, so the SSE streaming endpoint silently dropped the per-call thinking override entirely — a per-task mode wired only into the buffered route would have been invisible to any streaming UI. Both routes now honour both keys.
  • librefang-api: reasoning_mode on MessageRequest; both send_message and send_message_stream build the override with ThinkingOverride::resolve.
  • librefang-api: redacted_config_json now emits thinking.reasoning_mode, so GET /api/config exposes what POST /api/config/set accepts. The every_writable_config_leaf_is_readable parity guard (GET /api/config omits config fields that PUT accepts — write-only settings are invisible in the dashboard #6596) caught the omission — without it the dashboard would render the field blank and read a successful save back as "not configured".
  • Regenerated crates/librefang-api/tests/fixtures/kernel_config_schema.golden.json — the new ReasoningMode definition plus the ThinkingConfig property.
  • Docs: new docs/architecture/reasoning-mode-resolution.md; [thinking] sections updated in docs/src/app/configuration/page.mdx, docs/src/app/configuration/features/page.mdx and both zh mirrors; librefang.toml.example (with its xtask/baselines/config.sha256 recomputed).
  • Changelog fragment changelog.d/added/7946-per-agent-reasoning-mode.md.

Verification

Scoped to the four crates this PR touches. cargo clippy --workspace --all-targets was NOT run: the shared host ran out of disk during it (No space left on device while writing librefang-api's query cache), and the coordinator asked that it not be retried. CI runs the workspace-wide lane.

cargo check --workspace --lib                                        # clean
cargo clippy -p librefang-types       --all-targets -- -D warnings   # zero warnings
cargo clippy -p librefang-llm-drivers --all-targets -- -D warnings   # zero warnings
cargo clippy -p librefang-kernel      --all-targets -- -D warnings   # zero warnings
cargo clippy -p librefang-api         --all-targets -- -D warnings   # zero warnings
cargo test -p librefang-types                                        # 1031 + 5 + 49 + 3 + 4 + 2 passed, 0 failed
cargo test -p librefang-llm-drivers                                  # 657 lib + every integration bin passed, 0 failed
cargo test -p librefang-kernel                                       # KERNEL_TEST_RESULT
cargo test -p librefang-api --lib                                    # 1145 passed, 0 failed
cargo test -p librefang-api --test agents_routes_integration         # 72 passed, 0 failed
cargo test -p librefang-api --test config_schema_golden              # 1 passed (golden matches)
cargo test -p librefang-api --test openapi_spec_test                 # 2 passed (openapi.json regenerated, in sync)
cargo test -p librefang-api --test dead_route_audit_test             # 1 passed
cargo test -p librefang-api --test openapi_path_coverage_test        # 1 passed
python3 scripts/codegen-sdks.py                                      # SDKs regenerated, no diff
python3 scripts/check-changelog-attribution.py                        # OK

New tests:

Wire shape — crates/librefang-llm-drivers/src/drivers/openai.rs, each asserting the serialized body and what is absent as well as what is present:

  • deepseek_v4_direct_none_mode_disables_thinking_on_the_wire
  • deepseek_v4_direct_graded_modes_enable_thinking_with_effort (low / high / max)
  • deepseek_v4_without_a_mode_keeps_the_pre_7946_wire_shape
  • openrouter_routed_deepseek_uses_the_nested_reasoning_object (all four modes, max → xhigh)
  • openrouter_falls_back_to_the_budget_bucket_without_a_mode
  • generic_openai_compatible_model_maps_modes_to_top_level_effort
  • explicit_mode_wins_over_the_budget_bucket
  • kimi_disable_wins_over_an_explicit_reasoning_mode
  • a_rejected_reasoning_control_suppresses_both_spellings
  • deepseek_r1_gets_no_reasoning_control_in_either_dialect
  • no_payload_ever_carries_both_reasoning_controls — the exhaustive mutual-exclusion guard

Resolution order — crates/librefang-kernel/src/kernel/tests.rs:

  • test_reasoning_mode_global_reaches_an_agent_that_declares_nothing
  • test_reasoning_mode_per_agent_beats_global
  • test_reasoning_mode_per_call_beats_per_agent_beats_global
  • test_reasoning_mode_per_call_creates_thinking_config_when_absent
  • test_disable_and_mode_none_are_not_the_same_thing

Config / manifest surface — librefang-types:

  • test_thinking_config_parses_reasoning_mode_from_toml, test_thinking_config_reasoning_mode_defaults_to_none, test_thinking_config_rejects_an_unknown_reasoning_mode, test_reasoning_mode_serde_spelling_matches_as_str, test_thinking_override_resolve_precedence
  • test_manifest_thinking_reasoning_mode_round_trips_through_toml, test_manifest_thinking_without_reasoning_mode_stays_absent

API — crates/librefang-api/tests/agents_routes_integration.rs, #[tokio::test] against the production router:

  • test_reasoning_mode_field_accepted_by_message_endpoint
  • test_invalid_reasoning_mode_is_rejected — the conclusive one. MessageRequest does not deny unknown fields, so "accepted" alone would also pass for a typo'd key; a "banana" mode can only 422 if serde is genuinely parsing the field into ReasoningMode.
  • test_reasoning_mode_omitted_and_legacy_thinking_boolean_still_work
  • test_reasoning_mode_accepted_by_streaming_message_endpoint
  • test_per_call_reasoning_mode_wins_over_the_legacy_boolean — asserts the override the handler builds, i.e. the injection site, per CLAUDE.md's warning that an Option::None default compiles silently while disabling the feature

Prompt-ordering determinism (#3298) does not apply: nothing added here reaches a prompt. The one map-shaped surface in the outbound body, extra_body, was already a BTreeMap, and the new reasoning field is a fixed-shape single-key object.

Human-only verification (live daemon + real provider)

Not run here — live daemon plus a real LLM is human-only per CLAUDE.md. The wire shape is pinned by unit tests, but only a real endpoint proves the provider accepts the payload.

Start the daemon the usual way (librefang + the start subcommand) with a real key configured, then:

DeepSeek V4 direct — non-think, then max:

curl -sS -X POST localhost:4545/api/agents/$AGENT/message \
  -H 'content-type: application/json' -H "authorization: Bearer $LIBREFANG_TOKEN" \
  -d '{"message":"What is 2+2? One word.","reasoning_mode":"none"}' | jq .
# expect: 200, no thinking trace, output_tokens close to the answer length

curl -sS -X POST localhost:4545/api/agents/$AGENT/message \
  -H 'content-type: application/json' -H "authorization: Bearer $LIBREFANG_TOKEN" \
  -d '{"message":"Plan a three-stage migration.","reasoning_mode":"max"}' | jq .
# expect: 200, a visible thinking trace, materially more output tokens than the run above

To confirm the bytes rather than the outcome, run the daemon with RUST_LOG=librefang_llm_drivers=trace and grep the request bodies for thinking / reasoning:

  • none → "thinking":{"type":"disabled"} and no reasoning_effort
  • max → "thinking":{"type":"enabled"} and "reasoning_effort":"max"

OpenRouter-routed — the mutual-exclusion case, which is the one that 400s if it regresses. Point an agent at openrouter/deepseek/deepseek-v4-pro with a real OPENROUTER_API_KEY:

for m in none low high max; do
  curl -sS -o /dev/null -w "$m -> %{http_code}\n" \
    -X POST localhost:4545/api/agents/$AGENT_OR/message \
    -H 'content-type: application/json' -H "authorization: Bearer $LIBREFANG_TOKEN" \
    -d "{\"message\":\"ping\",\"reasoning_mode\":\"$m\"}"
done
# expect 4x 200. A 400 means the body carried both reasoning_effort and reasoning.effort.
# On the wire, max must appear as "reasoning":{"effort":"xhigh"}.

Per-agent and global rungs — put [thinking] reasoning_mode = "none" in the agent's agent.toml and [thinking] reasoning_mode = "max" in config.toml, then:

curl -sS -X POST localhost:4545/api/config/reload -H "authorization: Bearer $LIBREFANG_TOKEN"
curl -sS -X POST localhost:4545/api/agents/$AGENT/message \
  -H 'content-type: application/json' -H "authorization: Bearer $LIBREFANG_TOKEN" \
  -d '{"message":"ping"}'
# expect the agent's "none" to win over the global "max"

Streaming route — the same body against /api/agents/{id}/message/stream, to confirm the per-call override now reaches the SSE path (it did not before this PR).

Deferred / out of scope

  • Dashboard and TUI exposure — optional per the issue, omitted deliberately (above).
  • Anthropic / Gemini / Ollama drivers still read budget_tokens and ignore reasoning_mode. Their wire control is a token budget rather than an effort enum, so honouring the mode there means inventing a mode→budget mapping that would silently change the budget existing Claude users already configured and pay for. That is a different provider family and a different behavioural risk than the issue describes, and it wants its own decision rather than a mapping picked here — flagging rather than deciding unilaterally, per CLAUDE.md. Stated explicitly in the architecture page's "Providers that ignore the mode" table so it is a documented boundary, not a surprise.
  • The WebSocket chat path still carries the /think boolean only. Adding a mode there means changing the WS message envelope, which is the UI work above.
  • openapi.json and both xtask/baselines/*.sha256 files are regenerated and committed here (cargo xtask codegen --openapi's underlying test, then shasum -a 256 — the same bytes schema-check gen writes). scripts/codegen-sdks.py produced no SDK diff. CI's openapi-drift lane should therefore find everything in sync.
  • No dashboard or TUI change at all, including no i18n label. ConfigPage.tsx renders schema-driven config fields through t(config.fld_${key}, fieldLabelFallback(key)), so [thinking] reasoning_mode degrades to a humanized "Reasoning Mode" label rather than breaking; adding a key to en.json alone would trip locale-parity.test.ts.
  • ⚠️ One dashboard gap a maintainer should decide on, because it can lose a setting rather than merely fail to show it. dashboard/src/lib/agentManifest.ts lists thinking in FORM_TOP_LEVEL_KEYS, so serializeManifestForm rebuilds the entire [thinking] table from its form model (agentManifest.ts:78-82 type, :537-542 write, :999-1003 parse). A key the form model does not hold is a key a dashboard agent-edit silently deletes from agent.toml — so an agent pinned to reasoning_mode = "none" would quietly revert to the provider default of thinking-on, and be billed for the reasoning. The fix is ~6 lines of preservation (carry it as an opaque string through parse and write, no UI), not the optional form field. I wrote it, then dropped it on instruction to keep this PR to the config-plus-wire surface, and because dashboard/node_modules is absent on this host so the vitest lane could not be run locally. Worth doing in the same cycle as the UI work, ahead of anyone setting the key and then editing that agent in the dashboard.

houko added 2 commits August 31, 2026 18:58
Add `reasoning_mode = "none" | "low" | "high" | "max"` as a first-class
setting on the existing `[thinking]` table, so it is settable globally in
`config.toml`, per agent in `agent.toml`, and per task on both message
endpoints. Resolution order is per-call > per-agent > global > compiled
default, matching the shape already documented for `max_history_messages`
and `session_mode`.

Before this, the choice could not be expressed at all, and for DeepSeek
V4 — direct or via OpenRouter — the driver sent no reasoning fields
whatsoever, so every request ran at the provider default of thinking-on
at `high`. Prompt-level workarounds are not a substitute: the reasoning
tokens are still generated, still consumed from the output budget, and
still billed.

The mode is translated per provider in one resolver,
`reasoning_wire_fields`, keyed off `base_url` rather than the model id:

  * DeepSeek V4 direct — `thinking: {"type": "disabled"}` for `none`,
    otherwise `thinking: {"type": "enabled"}` plus `reasoning_effort`
    (`low` / `high` / `max`).
  * OpenRouter-routed — the nested `reasoning: {effort}` form, with
    `max` as `xhigh`. Never together with top-level `reasoning_effort`,
    which OpenRouter answers with HTTP 400.
  * Other OpenAI-compatible endpoints — top-level `reasoning_effort`,
    with `none` expressed by omission and `max` clamped to `high`.

`budget_tokens` remains the fallback bucket, and an unset
`reasoning_mode` leaves every wire shape byte-for-byte unchanged.

Also fixes the SSE streaming endpoint, which passed a hardcoded `None`
into the kernel and so dropped the per-call thinking override entirely,
and extends the #7769 strip-and-retry to clear both spellings of the
reasoning control.

The per-call key is a new `reasoning_mode` field rather than a widened
`thinking` boolean: the boolean is in the wire contract of every existing
client, and `thinking: false` only omits the opt-in while
`reasoning_mode: "none"` sends the provider's explicit non-think toggle.

Dashboard and TUI exposure is deliberately left out.

Refs #7946
An extra_body.reasoning_effort override on an OpenRouter-routed request
left the nested `reasoning` object reasoning_wire_fields had already set,
so the merged body could carry both spellings — the exact combination
OpenRouter answers HTTP 400 to. merge_extra_body now clears whichever
spelling an override key replaces.

Also fills in the changelog fragment's literal (#PR) placeholder with the
actual PR number.

@houko houko left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Automated review pass against CLAUDE.md's rules.

Pushed two mechanical fixes directly to this branch (commit 78e19a7):

  • merge_extra_body didn't clear the counterpart reasoning key on an extra_body override, so an OpenRouter-routed request with extra_params.reasoning_effort set could end up carrying both reasoning and reasoning_effort on the wire — the exact combination this PR's own invariant says OpenRouter 400s on. Added a regression test (extra_body_reasoning_effort_override_clears_the_nested_openrouter_spelling) alongside the existing coverage.
  • The changelog fragment ended with a literal, unfilled (#PR) placeholder instead of (#8083) — per changelog.d/README.md this means the PR would keep its generated release-notes line and appear twice.

Two remaining items left as inline comments below since they need judgment calls I can't make confidently from here (real provider error text, and a design decision on shared helpers):

  1. The strip-and-retry rejection matcher may not recognize an OpenRouter gateway's rejection of the new nested reasoning field, which would silently break the #7769 self-healing retry for that dialect.
  2. A minor duplicated "is this OpenRouter" substring check between reasoning_wire_family() and shared_guard_provider().

cargo check -p librefang-llm-drivers --lib, cargo clippy -p librefang-llm-drivers --all-targets -- -D warnings, and cargo test -p librefang-llm-drivers --lib all pass with the fix.


Generated by Claude Code

if status == 400
&& oai_request.reasoning_effort.is_some()
&& (oai_request.reasoning_effort.is_some() || oai_request.reasoning.is_some())
&& crate::llm_driver::llm_errors::is_unsupported_reasoning_effort_error(&body)

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The comment above says "Both spellings are stripped (#7946)", but the gate itself (is_unsupported_reasoning_effort_error, in crates/librefang-llm-driver/src/llm_errors.rs, unmodified by this PR) still requires the literal substring "reasoning_effort" in the rejection body. For an OpenRouter-routed request that now sends the nested reasoning: {"effort": ...} object instead, a gateway rejecting that field is more likely to name it "reasoning" than "reasoning_effort" in its error text — in which case this check returns false, the field is never cleared, record_reasoning_effort_rejected (line below) never fires, and the #7769 negative cache never populates for that model. The agent would then fail identically on every subsequent turn instead of self-healing after one retry, unlike the sibling temperature/max_tokens strips just above.

Same gate is duplicated at line ~2112 for the streaming path.

Worth confirming what real OpenRouter-adjacent gateways actually say in a 400 body for the nested field, and widening is_unsupported_reasoning_effort_error (or adding a sibling matcher) to catch the bare "reasoning" rejection if so. Leaving this as a comment rather than a same-PR fix since the correct pattern depends on real provider error text.


Generated by Claude Code

/// `reasoning: {"effort": …}` form or top-level `reasoning_effort` is
/// accepted. `openrouter/deepseek/deepseek-v4-pro` and a direct
/// `deepseek-v4-pro` are the same model behind two different dialects.
fn reasoning_wire_family(&self) -> ReasoningWireFamily {

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Minor: reasoning_wire_family() re-implements shared_guard_provider()'s base_url.to_ascii_lowercase().contains("openrouter") check a few lines below, for a different purpose (wire dialect vs. rate-limit bucket tag). Not a bug today, but a future host alias added to one and not the other would silently desync which dialect a request gets from which rate-limit bucket it's grouped under. Might be worth factoring the "is this OpenRouter" test into one shared helper both call — not blocking.


Generated by Claude Code

@cloudflare-workers-and-pages

cloudflare-workers-and-pages Bot commented Aug 31, 2026 •

Copy link
Copy Markdown

Deploying librefang-docs with  Cloudflare Pages  Cloudflare Pages

Latest commit: 4943d18
Status: ✅  Deploy successful!
Preview URL: https://8616e155.librefang-docs.pages.dev
Branch Preview URL: https://feat-7946-per-agent-reasonin.librefang-docs.pages.dev

View logs

`GET /api/config` omitted `thinking.reasoning_mode` while
`POST /api/config/set` accepted it, which the `every_writable_config_leaf_is_readable`
parity guard (#6596) turns into a test failure — and which would have made
the dashboard render the field blank and read a successful save back as
"not configured". `redacted_config_json` now emits it, through
`ReasoningMode`'s serde form rather than `Debug` so the value matches the
options the dashboard's dropdown offers.

DeepSeek R1 routed through OpenRouter would have started receiving a
budget-derived `reasoning.effort` it never received before, because the
OpenRouter family branch ran ahead of the `Strip` echo-policy check. R1
reasons unconditionally and exposes no wire toggle, so the field changes
the payload without changing the model. The `Strip` check now
short-circuits first, in both dialects.

Also regenerates the `KernelConfig` JSON Schema golden fixture for the
new `ReasoningMode` definition and the `ThinkingConfig` property.

Refs #7946
@houko

houko commented Aug 31, 2026

Copy link
Copy Markdown
Contributor Author

Two follow-up commits are now on the branch (tip 8a943b1f3), and one note on how they got there.

78e19a72c — extra_body override could still send both reasoning spellings

reasoning_wire_fields guarantees reasoning_effort and nested reasoning are mutually exclusive in its own output, but an extra_body override of one spelling was merged on top without clearing the other, so an OpenRouter-routed request could still carry both — the exact combination OpenRouter answers 400 to.
Pinned by extra_body_reasoning_effort_override_clears_the_nested_openrouter_spelling.

8a943b1f3 — read-path parity, and a DeepSeek R1 regression

GET /api/config omitted thinking.reasoning_mode while POST /api/config/set accepted it. That is a every_writable_config_leaf_is_readable parity-guard failure (#6596), and in the dashboard it would have rendered the field blank and read a successful save back as "not configured". redacted_config_json now emits it through ReasoningMode's serde form rather than Debug, so the value matches the options the dropdown offers.

Separately: DeepSeek R1 routed through OpenRouter would have started receiving a budget-derived reasoning.effort it never received before, because the OpenRouter family branch ran ahead of the Strip echo-policy check. R1 reasons unconditionally and exposes no wire toggle, so that field changes the payload without changing the model. The Strip check now short-circuits first, in both dialects.

Also regenerates the KernelConfig JSON Schema golden fixture for the new ReasoningMode definition and the ThinkingConfig property.

Note on history

The branch briefly diverged: 78e19a72c was pushed while 8a943b1f3 existed only locally on the same base, so the two were siblings rather than a chain. I rebased the local commit onto the pushed one rather than force-pushing over it — a force-push here would have silently dropped the extra_body fix and its test. The rebase applied cleanly and both changes are present in the tip; verified by grepping the merged tree for each fix before pushing.

Verification of the merged tip

cargo test -p librefang-llm-drivers --lib — 659 passed, 0 failed, run against 8a943b1f3 after the rebase (the earlier green was measured on the pre-rebase history, so it did not cover this combination).

Not re-run by me on the merged tip: cargo test -p librefang-api and the workspace-wide clippy. The workspace-wide clippy was already being skipped deliberately on this branch for disk reasons (the host hit 100% full during this batch), and CI covers both. If every_writable_config_leaf_is_readable is the guard that matters most to a reviewer here, that is the one to watch in the librefang-api lane.

@github-actions github-actions Bot added size/XL 1000+ lines changed area/docs Documentation and guides area/kernel Core kernel (scheduling, RBAC, workflows) labels Aug 31, 2026
houko and others added 3 commits August 31, 2026 05:31
…#8083

`reasoning_mode = "none"` turned reasoning ON for the budget-based drivers.
`apply_thinking_override` has to materialise a `ThinkingConfig` to carry the mode, and its default `budget_tokens` is 10_000.
Anthropic read that as "enable extended thinking" and inflated `max_tokens` by `budget + 1024`; Ollama's `think: Some(request.thinking.is_some())` read it as `think: true`.
An agent that asked not to reason reasoned, and was billed for it.
Both drivers now treat `reasoning_mode = "none"` as off.

The strip-and-retry never fired for the nested OpenRouter spelling.
`is_unsupported_reasoning_effort_error` still required the literal `reasoning_effort` in the body, so a gateway answering `unsupported parameter: reasoning` matched nothing: no strip, no retry, no negative-cache entry, and the same 400 on every subsequent turn.
Both names are recognised now; `reasoning_content` stays excluded, because DeepSeek's echo-it-back demand is a content requirement that stripping a control would not answer.

`thinking: true` was a silent no-op against an inherited `reasoning_mode = "none"`.
The boolean documents itself as "force thinking on even if the manifest has it off", and a non-think mode is exactly that off-state, so `Enable` now clears it. A graded mode is left alone.

An `extra_body` supplying both spellings ended up with neither.
Each guard removed the counterpart because the other key was present, so an operator who overrode both lost both and the model ran at its provider default with nothing in the payload to explain why.

Left as-is deliberately: the `effort_rejected` short-circuit in `reasoning_wire_fields` returns ahead of the `Echo` branch, so a cached rejection also suppresses V4's `thinking: {"type": "disabled"}` gate.
Emitting it anyway risks a permanent 400 on a gateway that does not understand `thinking` either — the retry loop clears only the two effort fields — so the conservative direction was kept and documented.

Verification: cargo check --workspace --lib; cargo clippy -p librefang-llm-driver -p librefang-llm-drivers --all-targets -- -D warnings; cargo clippy -p librefang-kernel --lib -- -D warnings; cargo test -p librefang-llm-driver --lib (68 passed); cargo test -p librefang-llm-drivers --lib (670 passed); cargo test -p librefang-kernel --lib (1720 passed).
The visual agent editor re-emits `[thinking]` from the two fields it has widgets for, and #7946 adds a third.
`parseManifestToml` read `budget_tokens` and `stream_thinking` into the form and dropped everything else on the floor, so opening any agent in the editor and pressing save rewrote its `agent.toml` without `reasoning_mode` — the per-agent knob the feature exists to provide, destroyed by the surface most operators will use to set it.
`[capabilities]`, `[model]` and `[resources]` already guard this with a `stripKnown` / extras round-trip; `[thinking]` had no such slot because until now the form was a complete description of the table.
It gets one, populated on parse and re-emitted inside the section on serialize, and dropped along with the rest when the user unticks the section.

`generateManifestMarkdown` carried an inline `ManifestExtras` literal that silently drifts every time the interface grows; it now calls `emptyManifestExtras()` so the compiler catches the next one, which is how this was caught.

The four configuration pages still said the Anthropic / Gemini / Ollama drivers "continue to use `budget_tokens`".
That stopped being true when those drivers learned to honour `reasoning_mode = "none"`, and a reader trusting the old sentence would conclude the non-think rung does nothing on Claude.
Corrected in both the English pages and their `zh` mirrors, matching the wording already in `docs/architecture/reasoning-mode-resolution.md`.

No changelog fragment: the loss is only reachable through #7946's own field, which has not shipped, so there is nothing for a release reader to have been bitten by.

Verification: `pnpm typecheck` clean, `pnpm eslint` clean on the changed files, `pnpm vitest run` 1674 passed across 200 files.
`round_trips an unknown [thinking] key such as reasoning_mode` was confirmed to fail against the unmodified source before the fix was applied.

Refs #7946
@houko

houko commented Sep 1, 2026

Copy link
Copy Markdown
Contributor Author

Review of #8083 — two follow-up commits pushed (b9985c6c6, 4943d187c)

Reviewed the full diff against origin/main, reading build_request, reasoning_wire_fields and merge_extra_body directly rather than inferring behaviour from the tests. All line numbers below are on the branch as it stands now.

The mutual exclusion holds — this part is right

reasoning_wire_fields (crates/librefang-llm-drivers/src/drivers/openai.rs:823) is structurally incapable of setting both spellings: the OpenRouter branch at :851 returns early with only reasoning, and the Echo / None arms below it are reachable only for Direct. no_payload_ever_carries_both_reasoning_controls (:4199) is not a spot check — it iterates 2 families × 4 policies × 5 modes × 5 budgets × 2 rejected states and asserts both the "never both" invariant and the per-family "only this spelling" invariant, so it covers the budget-derived fallback and the effort_rejected path as well as the explicit modes. Strip (DeepSeek R1) short-circuits at :843, ahead of the family branch, so R1 gets nothing in either dialect — including through OpenRouter, which is the case that would otherwise have regressed. The extra_body override path is guarded separately and its test asserts the absence of the counterpart key, not just the presence of the override.

Fixed in b9985c6c6

  1. crates/librefang-kernel/src/kernel/manifest_helpers.rs:219 — reasoning_mode = "none" turned reasoning ON for the budget-based drivers. ThinkingOverride::Mode has to materialise a ThinkingConfig to carry the mode, and ThinkingConfig::default().budget_tokens is 10_000. Anthropic (anthropic.rs:432, filter(|tc| tc.budget_tokens >= 1024)) and Ollama (ollama.rs:409, formerly think: Some(request.thinking.is_some())) steer on that budget alone, so none — per call, per agent, or global — enabled extended thinking and inflated Anthropic's max_tokens to budget + 1024. An agent asked not to reason, reasoned, and was billed for it. Both drivers now treat none as off, each with a positive control alongside.

  2. crates/librefang-llm-driver/src/llm_errors.rs:796 — the widened strip-and-retry could never fire for the nested spelling. The PR broadened the bug(llm): a gateway that rejects reasoning_effort fails every turn and then opens the provider circuit breaker #7769 retry gate to oai_request.reasoning.is_some(), but is_unsupported_reasoning_effort_error still required the literal string reasoning_effort in the body. A gateway answering Unsupported parameter: 'reasoning' matched nothing: no strip, no retry, no negative-cache entry, and the same 400 on every subsequent turn. Both names now match; reasoning_content stays excluded, since DeepSeek's echo-it-back demand is a content requirement that stripping a control would not answer.

  3. manifest_helpers.rs:209 — thinking: true was a silent no-op against an inherited none. Enable did nothing when manifest.thinking was already set, so the boolean whose own doc says "force thinking on even if the manifest has it off" lost to an agent-level or global reasoning_mode = "none". Enable now clears a none mode; graded modes are left alone.

  4. openai.rs:701 — an extra_body carrying both spellings ended up with neither. Each removal fired because the other key was present, so an operator who deliberately overrode both lost both and the model silently ran at its provider default. Each guard now also requires the counterpart to be absent.

Fixed in 4943d187c

  1. crates/librefang-api/dashboard/src/lib/agentManifest.ts:1017 — the visual agent editor deleted reasoning_mode on every save. parseManifestToml read budget_tokens and stream_thinking out of [thinking] and discarded the rest, and serializeManifestForm re-emitted the section from those two fields alone. Opening any agent in the dashboard editor and pressing save rewrote its agent.toml without reasoning_mode — the per-agent knob this PR exists to add, destroyed by the surface most operators will use to set it. [model], [resources] and [capabilities] already guard exactly this with a stripKnown → extras round-trip; [thinking] had no such slot because until this PR the form was a complete description of the table. It has one now, dropped along with the section when the user unticks it. The new test round-trips an unknown [thinking] key such as reasoning_mode was confirmed to fail against the unmodified source before the fix was applied.

  2. crates/librefang-api/dashboard/src/lib/agentManifestMarkdown.ts:57 carried an inline ManifestExtras literal that drifts silently every time the interface grows. It now calls emptyManifestExtras() so tsc catches the next one — which is how Bump tokio-tungstenite from 0.24.0 to 0.28.0 #5 was caught here.

  3. The four configuration pages were left stale by fix Bump docker/build-push-action from 6 to 7 #1. docs/src/app/configuration/page.mdx:1679 and its three siblings still said the Anthropic / Gemini / Ollama drivers "continue to use budget_tokens", which stopped being true once those drivers learned to honour none. A reader trusting that sentence would conclude the non-think rung does nothing on Claude. Corrected in both English pages and their zh mirrors, matching the wording already in docs/architecture/reasoning-mode-resolution.md.

Reported, not fixed

  1. openai.rs:847 — the effort_rejected short-circuit also suppresses V4's non-think gate. if effort_rejected { return out; } sits ahead of the Echo branch, so one 400 on a graded turn puts the model in the negative cache and a V4 agent pinned to reasoning_mode = "none" silently stops receiving thinking: {"type": "disabled"}, reverting to reasoning-on. Emitting the gate anyway risks a permanent 400 on a gateway that does not understand thinking either, since the retry loop clears only the two effort fields. The conservative direction was kept and the trade-off written into the architecture doc rather than decided silently.

  2. crates/librefang-api/src/routes/agents/messaging.rs:220 — the ephemeral: true / /btw branch calls send_message_ephemeral (:227), which takes no override, so reasoning_mode and the legacy boolean are accepted and silently discarded on that path. Pre-existing for thinking; closing it means a new parameter on a public kernel entry point, which is a wider blast radius than this diff.

Verified correct, no change needed

  • Resolution order is per-call > per-agent > global > compiled default, implemented in exactly one place. Global backfill (messaging.rs:2549, agent_execution.rs:678) fires only when manifest.thinking.is_none(), then apply_thinking_override runs last at both sites (:2560, :689) through the same shared helper — so the two call sites cannot disagree. The legacy Option<bool> reaches it through one From shim, and ThinkingOverride::resolve is the single place reasoning_mode beats the boolean.
  • Config-field mechanics are complete: struct field, #[serde(default)], Default entry, derives. redacted_config_json (routes/config/manage.rs:597) emits it on the global read path; the per-agent path serializes AgentManifest through serde, so it needs no hand-written entry. All three baselines recompute to their committed values — librefang.toml.example, openapi.json and examples/custom-agent/agent.toml each match xtask/baselines/*.sha256 byte for byte — and the schema golden carries both the ReasoningMode definition and the ThinkingConfig property.
  • bug(docs/config): [agents.<name>.proactive_memory] block in config.toml silently ignored — real path is agent.toml #5476: thinking was already in PER_AGENT_OVERRIDE_KEYS, so [agents.<name>.thinking] in config.toml warns rather than silently no-opping. Neither [thinking] block added to librefang.toml.example is nested under [agents.…].
  • Docs: the per-provider mapping table matches the code row by row — none → field omitted for generic OpenAI-compatible endpoints (generic_reasoning_effort returns None), max → xhigh on OpenRouter, max → "max" on V4 direct, max → "high" clamped elsewhere. The zh mirrors are faithful translations, not stale copies.
  • Rebase: git log --oneline origin/main..HEAD shows all three original commits plus the later main merge, and every change each commit describes is present in the tree.

Verification

Dashboard: pnpm typecheck clean, pnpm eslint clean on the changed files, pnpm vitest run — 1674 passed across 200 files. Rust (run on b9985c6c6, before the worktree target was cleaned for disk): cargo clippy -p librefang-llm-driver -p librefang-llm-drivers --all-targets -- -D warnings and cargo clippy -p librefang-kernel --lib -- -D warnings zero warnings; cargo test -p librefang-llm-driver --lib 68 passed, -p librefang-llm-drivers --lib 670 passed, -p librefang-kernel --lib 1720 passed. Eleven new regression tests across the two commits. No daemon was started and no live provider call was made — the DeepSeek / OpenRouter wire shapes are asserted against serialized bodies, not against a live endpoint, so that remains worth one human smoke test before merge.

@houko
houko merged commit dce46ae into main Sep 1, 2026
45 of 46 checks passed
@houko
houko deleted the feat/7946-per-agent-reasoning-mode branch September 1, 2026 07:18
@houko houko mentioned this pull request Sep 13, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/docs Documentation and guides area/kernel Core kernel (scheduling, RBAC, workflows) size/XL 1000+ lines changed

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Feature] Per-agent / per-task control of reasoning mode (non-think / low / high / max) for DeepSeek V4 & OpenRouter

1 participant