Decision up front
Documentation only. No code change. The race below is real and was verified, but what it leaves behind is leftover data — not a leak of memory, CPU, or behaviour — and the practice that avoids it is the same practice that is right for other reasons. So: describe it, and document the rule-writing guidance that makes it rare.
Reported by Codex in #149's final review round, never answered before merge, verified still present on main @ eceb676.
The race
A config reload quiesces every session manager (gateway/service.py:1311) and re-arms them much later (:1446). A membership removal — the bot being taken out of a room — that arrives inside that window is deferred rather than handled:
if self._watcher_manager.disarmed:
if self._quiesced:
self._deferred_membership.append(("removed", room_id)) # replayed on rearm
return
rearm() replays it. By then the reload has already stopped the old agent backends (stop_some, :1386) and installed the candidate's agent map (:1418). If the reload changed an agent's type or working_directory, or removed an agent, reclamation now finds a missing or mismatched backend identity and permanently skips deleting that room's backend session, prompt file and attachment workspace.
The window is not milliseconds: it spans stopping the old backends (stop_with_retries) and starting the new ones.
It also has a correlation that makes the independent estimate misleading — an operator reorganising agents is exactly the person likely to be adjusting which rooms the bot sits in during the same maintenance.
What it actually costs — the reason this is documentation and not a fix
Checked rather than assumed:
|
|
| CPU |
none |
| Gateway memory |
none |
| Scheduled jobs |
already cancelled — _on_membership_removed is documented as "reclaim its record, cancel its jobs", and that half runs |
| The watcher record |
removed correctly |
| Actually left behind |
the room's prompt file and attachment directory on disk, plus one session record inside the agent backend |
And the backend half is narrower still: only the OpenCode adapter implements delete_session (gateway/agents/opencode/adapter.py:1504, an HTTP DELETE /session/{id}). The base implementation returns False — deletion unsupported — so for a Claude backend nothing was ever going to be deleted and this defect changes nothing at all.
What to document
1. The race itself
A short note where reload behaviour is described (docs/design/config-reload-design.md, cross-referenced from docs/design/dynamic-watcher-design.md): a membership removal arriving during a reload that changes agent identity can leave a room's prompt file and attachment directory behind, and the operator-side avoidance is to take the bot out of the room first, let that settle, then reload. That ordering works because an un-quiesced manager handles the removal on the normal path, while the old backend is still alive, so the cleanup succeeds.
2. The rule-writing practice, in the user guide
This is the more valuable half, and it stands on its own merits — the race is just one more reason for it.
Write watcher rules with wide coverage, so day-to-day changes do not require a config reload or a daemon restart at all. Which rooms an agent is actually in belongs to the messaging platform: add or remove the bot there, and let the dynamic watcher machinery follow. config.yaml should describe what an agent is, not enumerate the rooms it may sit in.
A narrow rule set inverts that — every routine membership change becomes a config edit and a reload, which is both operational friction and the only way to be exposed to the race above.
Worth stating plainly in the guide: a deployment that needs a config reload to add or remove a room has not adopted dynamic watchers, and that is the signal to widen the rules rather than to keep editing them.
Place it near the watcher-rule material in docs/user-guide.md; docs/design/dynamic-watcher-design.md is the design-side reference to link.
Scope note
Widening the rules makes this rare, not impossible: the trigger is a reload that changes an agent's identity or removes an agent, and that is a genuine config.yaml change no amount of rule coverage eliminates. The documentation should say that rather than promise the practice closes it.
Decision up front
Documentation only. No code change. The race below is real and was verified, but what it leaves behind is leftover data — not a leak of memory, CPU, or behaviour — and the practice that avoids it is the same practice that is right for other reasons. So: describe it, and document the rule-writing guidance that makes it rare.
Reported by Codex in #149's final review round, never answered before merge, verified still present on
main@eceb676.The race
A
config reloadquiesces every session manager (gateway/service.py:1311) and re-arms them much later (:1446). A membership removal — the bot being taken out of a room — that arrives inside that window is deferred rather than handled:rearm()replays it. By then the reload has already stopped the old agent backends (stop_some,:1386) and installed the candidate's agent map (:1418). If the reload changed an agent'stypeorworking_directory, or removed an agent, reclamation now finds a missing or mismatched backend identity and permanently skips deleting that room's backend session, prompt file and attachment workspace.The window is not milliseconds: it spans stopping the old backends (
stop_with_retries) and starting the new ones.It also has a correlation that makes the independent estimate misleading — an operator reorganising agents is exactly the person likely to be adjusting which rooms the bot sits in during the same maintenance.
What it actually costs — the reason this is documentation and not a fix
Checked rather than assumed:
_on_membership_removedis documented as "reclaim its record, cancel its jobs", and that half runsAnd the backend half is narrower still: only the OpenCode adapter implements
delete_session(gateway/agents/opencode/adapter.py:1504, an HTTPDELETE /session/{id}). The base implementation returnsFalse— deletion unsupported — so for a Claude backend nothing was ever going to be deleted and this defect changes nothing at all.What to document
1. The race itself
A short note where reload behaviour is described (
docs/design/config-reload-design.md, cross-referenced fromdocs/design/dynamic-watcher-design.md): a membership removal arriving during a reload that changes agent identity can leave a room's prompt file and attachment directory behind, and the operator-side avoidance is to take the bot out of the room first, let that settle, then reload. That ordering works because an un-quiesced manager handles the removal on the normal path, while the old backend is still alive, so the cleanup succeeds.2. The rule-writing practice, in the user guide
This is the more valuable half, and it stands on its own merits — the race is just one more reason for it.
Write watcher rules with wide coverage, so day-to-day changes do not require a
config reloador a daemon restart at all. Which rooms an agent is actually in belongs to the messaging platform: add or remove the bot there, and let the dynamic watcher machinery follow.config.yamlshould describe what an agent is, not enumerate the rooms it may sit in.A narrow rule set inverts that — every routine membership change becomes a config edit and a reload, which is both operational friction and the only way to be exposed to the race above.
Worth stating plainly in the guide: a deployment that needs a
config reloadto add or remove a room has not adopted dynamic watchers, and that is the signal to widen the rules rather than to keep editing them.Place it near the watcher-rule material in
docs/user-guide.md;docs/design/dynamic-watcher-design.mdis the design-side reference to link.Scope note
Widening the rules makes this rare, not impossible: the trigger is a reload that changes an agent's identity or removes an agent, and that is a genuine
config.yamlchange no amount of rule coverage eliminates. The documentation should say that rather than promise the practice closes it.