Skip to content

bug: Telegram slash commands appear inert — SDK drift + command-path routing divergence #7140

Description

@DaBlitzStein

Symptom

On Telegram, /new, /agents, and related commands return a confirmation phrase but the visible conversation does not change.

What was verified (correction of the initial report)

The command dispatcher is fully implemented in crates/librefang-channels/src/bridge.rs:6828 (handle_command), with real kernel calls: "new" → reset_channel_session → kernel.reset_session, "agents" → list_agents (inline keyboard), "agent" → find/spawn + store user default, plus start/help/status/model/usage/.... The sidecar parses bot_command → Content.command and the daemon deserializes it into ChannelContent::Command (sidecar.rs:1149). The chain is real end-to-end.

Root cause candidates (ranked)

1. Sidecar SDK drift (deployment issue, most likely trigger)

The Telegram sidecar runs a pip-installed librefang SDK. A live deployment showed SDK 2026.3.2201 (March) against a daemon of 2026.7.31 (August). The sidecar protocol has evolved in between (capabilities negotiation, sender-identity fields, content handling). A 4-month-old adapter paired with a current daemon can produce exactly this symptom class: messages that arrive but are misparsed or silently degraded.

Recommendation: the daemon should log a WARN when a sidecar reports a protocol/sdk version older than the daemon expects (today it does not), and describe_sidecar should surface the adapter version.

2. Command-path routing divergence (code-level, real)

/new resolves the target agent via resolve_for_command (bridge.rs:6845) which consults ONLY the router chain (binding → direct route → per-account user default → channel default → system default). The regular message path resolve_or_fallback (bridge.rs:3511) additionally consults: addressed agent, #5671 conversation override, #5323 sticky holder, per-peer binding, and the channel_bindings SQLite instance default.

When those differ (sticky holder points at agent B, router resolves to agent A), /new resets agent A and replies "Session reset for this telegram chat" — a real reset on the wrong agent, invisible to the user.

3. /agent <name> selection does not stick

/agent writes only the router store (bridge.rs:6913). set_conversation_binding (channel_binding_store.rs:123) has no production caller — the #5671 conversation-binding store is write-dead. A selection made via /agent can therefore be overridden by older, higher-precedence bindings on the next message.

4. Stubs with confirmation phrasing (by design, but misleading)

  • /think (channel_bridge.rs:1974): explicit stub — stores a preference, returns "This will take effect when supported by the model."
  • /goal (channel_bridge.rs:1339): start_goal_run is fire-and-forget; if the kernel self-handle is unset the run silently never starts (goal_lifecycle.rs:43) while the reply says "Goal created and started".

Suggested fixes

  1. Version the sidecar protocol: daemon logs a WARN on old sidecar SDK versions and describe_sidecar shows the version. Ship/install the matching SDK next to the daemon binary.
  2. Unify command resolution: resolve_for_command should use the same chain as resolve_or_fallback (sticky holders + conversation overrides first, router fallback second).
  3. Wire /agent to set_conversation_binding so the selection binds the active conversation, or document which store is authoritative.
  4. Make /think and /goal honest: report "not yet supported" for /think until implemented; await/verify start_goal_run before confirming.

Activity

  1. changed the title [-]bug: Telegram slash commands (/new, /agents, /reset...) return fake confirmations — no execution path[/-] [+]bug: Telegram slash commands appear inert — SDK drift + command-path routing divergence[/+] on Aug 13, 2026
  2. houko commented on Aug 23, 2026

    @houko
    Contributor

    Coverage audit before this closes, because has-pr is currently overstating it.

    What the two linked PRs actually cover

    #7159 (fix/telegram-commands-origin) touches crates/librefang-api/src/channel_bridge.rs, crates/librefang-channels/src/bridge.rs, crates/librefang-kernel/src/kernel/messaging.rs, crates/librefang-kernel/src/kernel_api.rs — candidate 2 (resolve_for_command mirroring the chat chain), candidate 3 (set_conversation_binding now has production callers, 9 hunks reference it, commit 18ef5fd), and the /think stub, which is still live on main at channel_bridge.rs:1921 with an unused _agent_id.

    #7701 (fix/new-reset-canonical-session) touches crates/librefang-api/src/channel_bridge.rs only — the session dimension (derived sid vs canonical session).

    Root cause 1 is untouched by both. Neither diff includes crates/librefang-channels/src/sidecar.rs, crates/librefang-channels/src/embedded_sdk.rs, crates/librefang-api/src/routes/sidecar_describe.rs, or anything under sdk/. #7159's own body says so: "SDK version drift (deployment concern, not code)".

    It is not purely a deployment concern — there are four code-level gaps that are exactly why the drift was silent:

    1. The version field is never populated. SidecarAdapter.protocol_version defaults to None in both SDKs (sdk/python/librefang/sidecar/runtime.py:110, sdk/rust/librefang-sidecar/src/runtime.rs:139-141) and no first-party adapter overrides it — sdk/python/librefang/sidecar/adapters/telegram.py:830 declares capabilities only. docs/architecture/sidecar-protocol.md:70-71 states the current value is 1. So the ready frame from the affected deployment carried protocol_version: null.
    2. No expected value exists to compare against. crates/librefang-channels/src/sidecar.rs:1031-1036 logs protocol_version at INFO on the "Sidecar adapter ready" line and does nothing with it; SidecarReadyParams.protocol_version at sidecar.rs:256-259 is commented "Reserved for skew diagnostics (logged, never enforced)". There is no SIDECAR_PROTOCOL_VERSION constant anywhere in the repo to diff it against.
    3. The SDK package version never reaches the daemon on any path. librefang.__version__ exists (sdk/python/librefang/__init__.py:35, currently 2026.8.19 per sdk/python/pyproject.toml:7) and is not on the wire. Schema.to_dict() (sdk/python/librefang/sidecar/protocol.py:524-530) emits name, display_name, description, fields and nothing else; the daemon-side SidecarSchema (crates/librefang-api/src/routes/sidecar_describe.rs:30-36) has no field to receive it. So the describe_sidecar half of the suggested fix has no counterpart on either side.
    4. The mechanism that lets a stale install win, quietly. pythonpath_with_embedded skips the binary-embedded SDK whenever has_real_sdk_installed(command) is true (crates/librefang-channels/src/embedded_sdk.rs:369-375), and that probe is python -c "import librefang.sidecar" — importability only, no version read (embedded_sdk.rs:302-315). A March pip install therefore shadows the August tree compiled into the daemon, and the only trace is a debug!. That is the reported 2026.3.2201-vs-2026.7.31 pairing, reproduced from the code path rather than the deployment.

    A fix for 1 is a one-liner per adapter; 2 and 3 are the actual work, and 4 is what turns "operator has a stale SDK" from invisible into a WARN at spawn.

    The blocker on #7159. It carries a standing CHANGES_REQUESTED (2026-08-22, tip 6e34904) whose substance is unresolved on the current tip: the /think preference is keyed by agent alone — KernelBridgeAdapter::thinking_pref_key(agent_id) -> agent_id.0.to_string() — while the ack reads "Extended thinking {state} for this chat." One agent serving several Telegram chats (or several channel instances) means /think on in one DM changes reasoning mode and cost for every other conversation routed to that agent, and /think off anywhere disables it for all of them. The pref needs the conversation identity in the key, threaded through set_thinking and both send paths, plus a regression with two conversations on one agent proving isolation. CI is otherwise green on that tip, so this is the only thing between #7159 and merge.

    Not blocking: the /goal half of candidate 4 is unreachable on main — there is no /goal command in the channel bridge, since #6505 was closed. It should not be counted against this issue.

    So: #7159 (after the scoping fix) + #7701 close candidates 2, 3, the /think stub and the session dimension. Root cause 1 — the headline half of this issue's title — has no PR.

  3. houko commented on Aug 24, 2026

    @houko
    Contributor

    Closing this as part of a backlog sweep. Re-verified against origin/main @ 45e9bf0. Two of the four root causes are closed and two are not, so this closes as not planned rather than completed.

    What landed

    Root cause 1 (SDK drift) — closed, and my earlier triage on this thread was out of date the moment #7848 merged. I wrote here that "root cause 1 — the headline half of this issue's title — has no PR". It has one, and it merged: #7848 (7bb78c9500707e0bf684e79738dc78169a240f7f), which closes all four of the code-level sub-gaps I enumerated:

    • SIDECAR_PROTOCOL_VERSION now exists (crates/librefang-channels/src/sidecar.rs:271), with classify_protocol_version at :291 turning a declared version into Match / Older / Unspecified, and the ready handler at sidecar.rs:1068 emitting a WARN on skew and on an absent value instead of an unexamined INFO line.
    • The SDKs now populate it: sdk/python/librefang/sidecar/runtime.py:113 defaults protocol_version to protocol.PROTOCOL_VERSION, and the Rust SDK does the same at sdk/rust/librefang-sidecar/src/runtime.rs:142.
    • The SDK package version reaches the daemon: sdk/python/librefang/sidecar/protocol.py:551 emits sdk_version unconditionally, and crates/librefang-api/src/routes/sidecar_describe.rs:40 has a field to receive it.
    • The stale-install path is no longer silent: installed_sdk_version (crates/librefang-channels/src/embedded_sdk.rs:334) reads the installed version rather than probing importability alone, and :413 compares it against embedded_sdk_version() and warns when they differ. That is exactly the 2026.3.2201-vs-2026.7.31 pairing you reported, now visible at spawn.

    Your recommendation in the issue body — that the daemon should warn when a sidecar reports an older protocol/SDK version, and that describe_sidecar should surface the adapter version — is what shipped, in both halves.

    Root cause 4, /think half — closed. #7854 (f5d2a5bd40ed7f9eab9136d3fe964bcc798e7c71). /think is no longer a stub that stores a preference and promises future support: it is keyed by conversation rather than by agent, and it is read on the send path. Write side crates/librefang-channels/src/bridge.rs:7094; key derivation and read side crates/librefang-api/src/channel_bridge.rs:646 with thinking_override_for at :647. The isolation regression proving one chat's toggle does not rewrite another chat's turns is at channel_bridge.rs:3425, and the command-path regression is at bridge.rs:8329. model_rejects_thinking (channel_bridge.rs:654) also stops the ack from promising extended thinking on a model whose catalog entry says it has none.

    Root cause 4, /goal half — not applicable. There is no /goal command in the channel bridge on main; grepping for it in bridge.rs and channel_bridge.rs returns nothing. It should not be counted against this issue.

    What remains

    Root cause 2 (command-path routing divergence) — open, exactly as you described it. resolve_for_command (crates/librefang-channels/src/bridge.rs:6872) still resolves through router.resolve_with_context and nothing else. The regular message path, resolve_or_fallback (bridge.rs:3531), additionally consults the thread-route agent (:3541), the #5671 conversation binding and the channel-instance binding (:3576), and the #5323 sticky holder and explicit @-addressing (:3584). When those disagree, /new at bridge.rs:6973 still resets the router's agent and replies as if it reset the one the user is talking to. Your ranking of this as a real code-level defect holds.

    Root cause 3 (/agent does not stick) — open. set_conversation_binding (crates/librefang-memory/src/channel_binding_store.rs:123) still has zero production callers; every reference on main is a test (channel_binding_store.rs:275, :299, :319, :322, and crates/librefang-api/src/channel_bridge.rs:3572, which sits inside the #[cfg(test)] block starting at :3147). The /agent arm at bridge.rs:6940-6947 writes only router.set_user_default_for_channel / set_user_default. The write-dead store you identified is still write-dead.

    Where the work lives

    PR #7159 — open, but not in a landable state: CHANGES_REQUESTED, mergeable: CONFLICTING / DIRTY, last pushed 2026-08-22 (6e349043689f057f873a70fc6b74a5f2c961dd12), and currently red on Quality, OpenAPI Drift, Test / Unit (lib+bin), Test / Ubuntu shard 1/4, Build / Linux aarch64 and Kernel Ignored / Restart Override. Its /think half has since been superseded by #7854, so what is left in it is the root cause 2 and 3 work, on a base several days stale.

    PR #7701 — open, single file crates/librefang-api/src/channel_bridge.rs, MERGEABLE but BLOCKED, red on Quality, Test / Unit (lib+bin), Test / Ubuntu shard 1/4 and Build / Linux aarch64. It covers the session dimension (/new, /reboot and /compact resetting the canonical session), not the routing divergence.

    Neither PR is closed by this sweep; closing the issue does not touch them.

    Closing note

    This is a backlog sweep, not a decision that root causes 2 and 3 are unwanted — they are confirmed bugs and I have said so twice on this thread. This was a good report, and the correction you supplied in it, that the dispatcher is fully implemented contrary to the initial framing, is what made the rest of it tractable. Reopening this, or filing a narrower issue for the resolve_for_command divergence alone now that root causes 1 and 4 are out of the way, is welcome.

  4. reopened this on Aug 24, 2026
  5. houko commented on Aug 24, 2026

    @houko
    Contributor

    Reopened — closing this as not planned was a mistake on my part.

    not planned reads as "we do not intend to do this", which is not true of this issue.
    The work is either in flight in an open pull request or tracked as real remaining work, and the preceding comment says which.

    That comment's content stands: what landed, what remains, and where the work lives are all accurate.
    Only the closure was wrong, and the label it carried misrepresented the state to anyone reading the tracker.

    Issues covered by an open PR will close on their own when that PR merges, with the correct reason.

  6. removed
    has-prA pull request has been linked to this issue
    on Aug 26, 2026
  7. houko commented on Sep 1, 2026

    @houko
    Contributor

    Not on main, so not closing — but the symptom decomposes into two separate mechanisms, and each now has somewhere to live. Posting because this has sat under awaiting-response since 2026-08-13 and both halves have been diagnosed since, elsewhere.

    /new returning a confirmation while the conversation is unchanged — #7701.

    That is exactly the mechanism, and this issue's own correction ("the command dispatcher is fully implemented … the chain is real end-to-end") is why it was hard to see: the dispatcher runs, the kernel call happens, the ack is honest — it just resets a different session from the one the user is looking at. #7701's changelog puts it as "a reset used to ack success against a session holding no messages while the conversation the user was looking at kept its full history", with a live case where the derived session held 0 messages and the canonical one held 198.

    So neither of this issue's ranked root-cause candidates is needed to explain the /new half: it is not SDK drift and not a routing divergence in the dispatcher, it is which SessionId the reset targets.

    Worth reading my review on #7701 before treating it as done — as written it resets both sessions, which takes the WebUI conversation with it while the ack still says "Other surfaces untouched".

    /agents inline keyboard not switching agent — #8118.

    Different mechanism, separately reported, with #8117 open against it. My review there found #8117's stated premise does not hold (dispatch_message already routes slash-prefixed button actions at bridge.rs:4538), and suggested the likely real cause: content_to_text's ButtonCallback arm at bridge.rs:1056 renders [Button: /agent X] with no slash passthrough, and it feeds the debounce/coalesce path — so a button pressed inside a debounce window is dispatched as prose and never reaches the command dispatcher.

    Suggestion: close this one in favour of #7701 and #8118 once either lands, or re-scope it to whichever half turns out not to be covered. The SDK-drift candidate ranked #1 here is worth keeping on record either way — it would produce the same surface symptom for a different reason, and nothing in the two PRs above addresses a stale pip-installed librefang SDK.

    Verified against main at 07cb6fc76.

  8. removed
    has-prA pull request has been linked to this issue
    awaiting-responseMaintainer replied — waiting for reporter/contributor feedback
    on Sep 23, 2026
  9. added a commit that references this issue on Sep 23, 2026
    8a4fe05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area/channelsMessaging channel adaptersarea/sdkJavaScript and Python SDKs

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions