Skip to content

RFC: add managed configuration and declarative provisioning mode #6695

Description

@whatnick

Summary

LibreFang currently treats $LIBREFANG_HOME/config.toml as writable application state.
This works for local and interactive deployments, but it does not provide a clean infrastructure-as-code contract for Kubernetes operators who want configuration sourced from an immutable ConfigMap or another deployment-managed file.

Add a managed configuration mode that separates immutable operator-owned configuration from mutable runtime data, enforces the distinction in the API, and presents provisioned settings as read-only in the dashboard.

Motivation

The supported Kubernetes baseline stores both configuration and runtime state under /data.
An operator can copy a ConfigMap into the PVC during startup, but this creates ambiguous ownership: copying on every restart overwrites dashboard changes, while copying only on first boot allows configuration to drift away from the manifest.
Mounting the ConfigMap directly is also unsafe today because LibreFang expects config.toml to be writable and several API routes update it.

Grafana-style provisioning provides a clearer model: deployment-managed resources are visible in the UI, marked with their provenance, and cannot be edited through interactive surfaces.
LibreFang should provide the same explicit contract rather than relying on filesystem permissions or UI-only disabling.

The immediate use case is OAuth/OIDC configuration under [external_auth], where public provider metadata belongs in a ConfigMap and client secrets belong in Kubernetes Secrets.
The design should apply to the complete kernel configuration rather than introducing a one-off OAuth path.

Current behavior

  • config.toml is loaded from LIBREFANG_HOME and is expected to remain writable.
  • /api/config/set mutates the file and applies the config reload plan.
  • The dashboard exposes editable controls for writable configuration paths.
  • [external_auth] is already deliberately excluded from the generic writable-path allowlist, but there is no global managed-mode status or provenance model.
  • The Kubernetes base injects authentication and provider secrets through environment variables while retaining configuration on the writable /data PVC.
  • Config reload already classifies fields as hot-reloadable, restart-required, or read-live, which provides a useful foundation for applying externally managed updates.

Proposed design

1. Separate configuration source from runtime state

Introduce an explicit configuration path independent of LIBREFANG_HOME, for example:

LIBREFANG_CONFIG_PATH=/etc/librefang/config.toml
LIBREFANG_CONFIG_MODE=managed

LIBREFANG_HOME=/data would continue to own SQLite, agents, logs, credentials, and other mutable runtime state.
The managed configuration file could then be mounted read-only from a ConfigMap without copying it into the PVC.

The mode must be selected outside the managed file itself so an API write cannot disable its own lock.
The default must remain the current mutable behavior for compatibility with desktop, local, Compose, and existing Kubernetes installations.

2. Enforce managed mode server-side

When managed mode is active, every endpoint that persists deployment configuration must reject the write consistently, preferably with 423 Locked and a structured response such as:

{
  "error": "configuration is managed by the deployment",
  "code": "config_managed",
  "source": "/etc/librefang/config.toml"
}

The audit must include more than /api/config/set.
Password changes, quick initialization, provider setup, and any domain route that writes config.toml must either be locked or explicitly classified as mutable runtime state.
Filesystem read-only errors must never be the enforcement mechanism.

3. Expose configuration provenance

Expose authenticated metadata through the config API, either alongside the redacted config or through a dedicated status endpoint:

{
  "mode": "managed",
  "source": "/etc/librefang/config.toml",
  "writable": false,
  "checksum": "sha256:...",
  "loaded_at": "..."
}

Do not expose secret values or sensitive filesystem details to unauthenticated callers.
The checksum should cover the effective non-secret configuration or have clearly documented semantics.

4. Make the dashboard accurately read-only

The dashboard should display managed configuration normally, add a visible “Managed by deployment” badge, disable or remove write controls, and explain where changes must be made.
The UI must consume server-provided capability/provenance metadata rather than infer managed mode from failed writes.
Operational actions that do not mutate deployment configuration should remain available.

5. Define update and reload behavior

Document one supported Kubernetes update workflow.
A ConfigMap hash annotation can trigger a StatefulSet rollout, which is the simplest initial contract and works for restart-required fields.
If in-place reload is supported, define whether it is triggered by POST /api/config/reload, SIGHUP, filesystem watching, or a sidecar, and ensure updates are read atomically rather than observing ConfigMap symlink swaps mid-read.

The API should report the same hot/restart/read-live classification already described by the reload planner.
A failed managed reload must retain the last valid effective configuration and expose an actionable error rather than partially applying the new file.

6. Keep secrets outside ConfigMaps

Managed configuration should continue to reference secret environment-variable names or secret-file paths.
The Kubernetes example must place OAuth client secrets, API keys, the vault key, and LIBREFANG_STATE_SECRET in Secrets rather than serializing them into config.toml.

Example target deployment shape:

env:
  - name: LIBREFANG_HOME
    value: /data
  - name: LIBREFANG_CONFIG_PATH
    value: /etc/librefang/config.toml
  - name: LIBREFANG_CONFIG_MODE
    value: managed
volumeMounts:
  - name: config
    mountPath: /etc/librefang
    readOnly: true
  - name: data
    mountPath: /data

Declarative resource provisioning

The RFC should decide whether the first implementation covers only KernelConfig or establishes a reusable provenance model for later Grafana-style resource provisioning, for example:

/etc/librefang/provisioning/
  config.toml
  agents/*.toml
  channels/*.toml
  workflows/*.toml

A later resource reconciler would need stable identifiers, idempotent application, provenance, drift reporting, conflict handling, and an explicit deletion/prune policy.
Provisioned resources should be locked individually while runtime-created resources remain mutable.
This broader reconciler does not need to block the initial managed config.toml mode, but the initial API metadata should avoid making it impossible.

Security considerations

  • Enforce locks in API handlers, not only in the dashboard.
  • Do not allow the managed file to turn managed mode off.
  • Keep Kubernetes Secret material out of ConfigMaps, API responses, logs, and checksums.
  • Preserve the existing OAuth state-secret boot validation.
  • Continue applying the writable-path allowlist and authorization checks in mutable mode.
  • Treat symlink replacement and partial file writes as untrusted reload inputs and parse into a complete candidate before swapping effective configuration.

Non-goals

  • Active-active or multi-replica LibreFang.
  • A Kubernetes Operator or CRDs in the initial implementation.
  • Replacing SQLite or moving runtime state into ConfigMaps.
  • Making all runtime operations read-only.

Acceptance criteria

  • A new installation can mount config.toml read-only outside /data and boot without an init-container copy.
  • Mutable mode remains the default and preserves existing behavior.
  • Managed mode is selected through a deployment-controlled input that cannot be changed through LibreFang's config API.
  • Every API route that persists deployment configuration is identified and locked server-side in managed mode.
  • Locked writes return one documented status and structured error shape.
  • The authenticated API exposes managed-mode provenance and current config status without leaking secrets.
  • The dashboard clearly marks managed settings and does not offer write controls for them.
  • Operational actions and mutable runtime state remain usable.
  • Invalid externally supplied configuration never partially replaces the last valid effective configuration.
  • Kubernetes documentation includes a ConfigMap, Secret references, checksum-triggered rollout, OAuth/OIDC example, and rollback procedure.
  • Integration tests cover managed boot, rejected writes, UI capability metadata, reload failure, secret redaction, and mutable-mode compatibility.
  • The supported deployment remains a single StatefulSet replica.

Design questions

  • Should the public contract use LIBREFANG_CONFIG_MODE=managed, infer managed mode from LIBREFANG_CONFIG_PATH, or require both?
  • Should managed mode lock the entire kernel configuration or allow field-level ownership overlays?
  • Which existing write routes mutate deployment configuration versus legitimate runtime state?
  • Should the first release support rollout-only updates or also in-place reload?
  • Should provenance metadata live in GET /api/config, GET /api/config/status, or a general capabilities endpoint?
  • How should configuration supplied through include files report provenance and checksums?
  • What deletion semantics should later provisioned agents and workflows use?

Related work

Activity

  1. added
    needs-triageAuto-applied when the issue title/body matched no area label — needs maintainer review
    on Aug 1, 2026
  2. houko commented on Aug 4, 2026

    @houko
    Contributor

    This RFC describes current behaviour accurately, and the four design questions at the end are answerable from the code. Taking them in order, with the audit item first since it is also an acceptance criterion.

    "Which existing write routes mutate deployment configuration versus legitimate runtime state?"

    /api/config/set is not the only surface, and the gap is wider than the RFC assumes. Routes that persist into config.toml today, each with a test that pins the behaviour:

    • PUT /api/budget — budget_routes_test.rs:279 is literally named budget_put_persists_to_config_toml.
    • PUT /api/providers/{name}/budget — provider_budget_routes_test.rs:83, "persist seed config so PUT round-trips through a real config.toml".
    • User creation — users_test.rs:524, users_create_refuses_to_overwrite_corrupt_config_toml.
    • Channel removal — channel_remove_test.rs:3, "so the config.toml rewrite lands in the sandbox".
    • Provider, skill, extension, and memory routes all reach the same file (routes/{providers,skills/*,memory,backup,channels,users,config/*}.rs).

    And one that is not a route at all, which is the sharpest problem for a read-only mount: crates/librefang-kernel/src/config.rs:213 writes the migrated config back to disk at boot whenever migrated && file_version < CONFIG_VERSION. A ConfigMap mounted read-only makes that std::fs::write fail on the first boot after any config-schema bump. It currently degrades to a tracing::warn! rather than aborting, so it will not crash — but it will re-run the migration on every subsequent boot, silently, forever. Managed mode has to make this an explicit, once-per-boot "migration required, config is deployment-managed" error rather than a warning that scrolls past.

    The RFC's instinct that "filesystem read-only errors must never be the enforcement mechanism" is right, and this is the concrete case that proves it.

    "Should the public contract use LIBREFANG_CONFIG_MODE=managed, infer from LIBREFANG_CONFIG_PATH, or require both?"

    Both, kept orthogonal.

    Relocating the file and locking it are different wants. An operator may reasonably want config.toml outside LIBREFANG_HOME while still editing it from the dashboard — a Compose deployment with a bind-mounted config directory, say. If the path implies the lock, that operator gets a read-only dashboard they never asked for, and the only way out is moving the file back.

    So: LIBREFANG_CONFIG_PATH relocates and nothing more. LIBREFANG_CONFIG_MODE=managed locks, and is meaningful with or without a custom path. The RFC's own constraint — the mode must not be settable from inside the managed file — is satisfied either way, and keeping them independent means the Kubernetes manifest states its intent explicitly rather than encoding it in a path.

    "Should managed mode lock the entire kernel configuration or allow field-level ownership overlays?"

    Whole config, in the first implementation.

    This codebase already has a field-level write model: WRITABLE_EXACT_PATHS and WRITABLE_SECTION_PREFIXES in crates/librefang-api/src/routes/config/mod.rs:666. It is already load-bearing for security rather than convenience — external_auth. is deliberately absent from it because, per the comment at routes/config/manage.rs:158, flipping an endpoint or the verification gate post-auth is the #3703 impersonation vector. There is also a writable ⊆ readable parity guard built on top of it.

    Layering a second, orthogonal ownership axis over that produces a two-dimensional matrix — writable-by-allowlist × owned-by-deployment — where the interesting cases are the corners, and where a future contributor adding a field has to reason about both. The security value of the existing allowlist comes from it being the single answer to "can this be written over HTTP". Keep it that way: in managed mode the answer is no, uniformly.

    Field-level overlays are a reasonable v2 once someone has a deployment that actually needs one. Designing for it now costs clarity and buys nothing that has been asked for.

    "Should the first release support rollout-only updates or also in-place reload?"

    Rollout-only.

    The reload classification the RFC wants to reuse already exists and is already documented: build_reload_plan in crates/librefang-kernel/src/config_reload.rs decides hot-reload / restart-required / read-live, and docs/operations/config-reload.md is the ops-facing table, drift-guarded by a test. Exposing that classification through the API is cheap and worth doing in v1.

    Acting on it in place is not. In-place reload of a ConfigMap means handling the symlink-swap semantics of Kubernetes' atomic update — the RFC names this itself — plus deciding between POST /api/config/reload, SIGHUP, a watcher, and a sidecar, and then guaranteeing that a partially-written or invalid file never replaces the last valid effective configuration. That is a second hard problem stacked on the first.

    A checksum annotation triggering a StatefulSet rollout works for restart-required fields too, which is the superset. Ship that, expose the classification, and let in-place reload be justified by a deployment that finds the rollout too slow.

    One thing to add to the acceptance criteria

    The criteria cover rejected writes but not the boot path. Worth adding: a managed-mode boot with a config file requiring schema migration produces a single actionable error and does not retry the write on every restart. That is the case the current config.rs:213 warning would otherwise hide.

  3. github-actions commented on Aug 7, 2026

    @github-actions
    Contributor

    🤖 PR #6717 just merged and appears to be the last open PR referencing this umbrella issue.

    Dry-run flag from the umbrella-autoclose workflow. Close manually if correct, or reopen / add an unchecked task-list item if more work is planned. Delete this comment to let the bot re-flag later.

  4. houko commented on Aug 8, 2026

    @houko
    Contributor

    #6717 merged as 35d3ab5, so here is what it covers against this RFC and what is left. It intentionally said Refs rather than Closes, and I think this issue should stay open — four acceptance criteria and one design question are still outstanding.

    Covered.
    The two-env-var contract with relocation and locking kept orthogonal, which is design question 1 answered as "both, independent".
    Server-side enforcement through a single guard_config_write() (crates/librefang-api/src/routes/mod.rs:260) returning 423 Locked with the {ok:false, error, code:"config_managed", source} shape this RFC specified, refused before the handler reads the file — the reasoning for enforcing in-process rather than relying on a read-only mount is recorded at routes/mod.rs:255.
    GET /api/config/status provenance with mode, source, writable, checksum and modified_at.
    Whole-config lock rather than field-level overlays (design question 2) and rollout-only updates (design question 4).
    Plus one thing this RFC did not mention: a managed-mode load now skips the boot-path migration write-back and logs one targeted warning, instead of retrying a doomed std::fs::write on every boot forever.

    Still open.

    1. The route audit is incomplete, and wider than the PR's own Known-gaps section says.
      Three provider handlers still rewrite config.toml in managed mode: set_provider_key (crates/librefang-api/src/routes/providers.rs:1452) through persist_default_model at :1610 and :1649, set_provider_url (:2408) through upsert_provider_url (:2499) and upsert_provider_proxy_url (:2503), and set_default_provider (:2590) through persist_default_model (:2659), with the writes at :2769, :3016 and :3061.
      The Known-gaps list in docs/operations/managed-config.md names skills, memory and server.rs instead, and that list is wrong in both directions: routes/memory.rs and src/server.rs contain no atomic_write calls at all, routes/skills/mod.rs writes skill secrets through atomic_write_secret_file (:1052, :1073, :1085) rather than config.toml, and providers.rs — the one real gap — is not mentioned.
      fix(api): lock the provider config routes in managed mode and correct the documented gap list #6737 guards the three handlers and corrects the list.

    2. RFC section 4 (the "Managed by deployment" badge and disabled write controls in the dashboard) is deferred.

    3. RFC section 5 (in-place reload / ConfigMap watching) is deliberately out of scope in favour of rollout-only.

    4. Field-level ownership overlays are deliberately out.

    5. The Kubernetes acceptance criterion is met as prose only.
      deploy/kubernetes/ is untouched; docs/operations/managed-config.md mentions ConfigMap and the checksum annotation in sentences at lines 7 and 74, with no manifest, no Secret-reference example, no OAuth/OIDC example and no rollback procedure.
      Section 6's requirement that the example place OAuth client secrets, API keys, the vault key and LIBREFANG_STATE_SECRET in Secrets therefore has no artefact behind it yet.

    6. The design question about include files is unanswered.
      The checksum is computed over the primary file's raw bytes, so an edit to an included file leaves it unchanged and an operator using the checksum to confirm a rollout landed gets a false negative.

    Integration coverage is also partial: crates/librefang-api/tests/config_managed_mode_test.rs has three cases (status-mutable, status-managed, and config/set refused with the file left untouched), and there is nothing yet for the newly guarded budget and users routes, secret redaction, or a managed-mode reload failure.

    I will convert the six items above into a checklist in this issue body so close-umbrella-on-last-pr.yml does not flag it, and drop the stale needs-triage.

  5. houko commented on Aug 23, 2026

    @houko
    Contributor

    Re-audited this against main today. The write guard from my 2026-08-08 comment has held up and grown, but it satisfies roughly the first half of the RFC, so this should stay open. Details, then the one new defect I found while checking.

    Verified as shipped

    guard_config_write() lives at crates/librefang-api/src/routes/mod.rs:361 and returns exactly the shape section 2 asked for — 423 Locked with {ok:false, error, code:"config_managed", source} — reading the mode from the environment on every call rather than caching it at boot.

    Eight call sites, all before the handler touches the file:

    Handler Guard
    config_set (routes/config/manage.rs:1277) :1386
    quick_init (routes/config/system.rs:144) :190
    persist_budget (routes/budget.rs:592) :596
    persist_users (routes/users.rs:1176) :1184
    memory_config_patch (routes/memory.rs:1700) :1704
    set_provider_key (routes/providers.rs) :1481
    set_provider_url :2422
    set_default_provider :2630

    The last three are #6737 (5352cce5a), which closed the gap I named last time.
    Section 3 is shipped too: GET /api/config/status (routes/config/manage.rs:1143, registered at routes/config/mod.rs:42) returns ConfigProvenance with mode / source / writable / checksum / modified_at, and the managed-mode boot skips the migration write-back with one targeted warning (crates/librefang-kernel/src/config.rs:289).

    Not the whole RFC

    Four acceptance criteria and the second half of the title are still unmet, and none of them is blocked on the guard.

    Section 4, the dashboard, has no consumer at all. Grepping crates/librefang-api/dashboard/src for config/status, ConfigStatus, config_managed and 423 returns nothing. So the criterion "the dashboard clearly marks managed settings and does not offer write controls for them" is unmet, and the UI still learns about managed mode by attempting a save and reading the refusal back — precisely what the doc-comment on config_status says the endpoint exists to avoid.

    The Kubernetes criteria have no artefact. deploy/kubernetes/base/ is kustomization.yaml, service.yaml, statefulset.yaml — no ConfigMap, and the statefulset's env block sets LIBREFANG_HOME=/data with neither LIBREFANG_CONFIG_PATH nor LIBREFANG_CONFIG_MODE. Nothing under deploy/ mentions either variable. So "a new installation can mount config.toml read-only outside /data and boot without an init-container copy" is documented prose only, and section 6's requirement that the example place the OAuth client secrets, API keys, vault key and LIBREFANG_STATE_SECRET in Secrets has nothing behind it.

    Declarative provisioning — the second half of this issue's title — does not exist. No provisioning/ tree, no resource reconciler, no per-resource provenance or prune policy. The only provisioning paths in the repo are Grafana's under deploy/grafana/.

    The route audit is still incomplete by decision, not by oversight. docs/operations/managed-config.md:117-158 enumerates what still writes config.toml in managed mode — the sidecar-channel configure/delete pair, the MCP server routes under mcp_runtime_store = "file", the extension install/uninstall routes unconditionally, and change_password in server.rs — and says locking them is a product decision. That decision is still outstanding, and it is the criterion "every API route that persists deployment configuration is identified and locked server-side".

    Integration coverage is now six cases in crates/librefang-api/tests/config_managed_mode_test.rs (status mutable, status managed, config/set, memory patch, and the three provider routes). Still nothing for the guarded budget, users and quick-init paths, nothing for secret redaction, and nothing for a managed-mode reload failure.

    Two rows of that Known-gaps table are now stale in the other direction: PATCH /api/memory/config and POST /api/init are listed as gaps but are both guarded, since #6982 (c6098ae2c) and #6988 (9797b3e0e) respectively — both landed after 5352cce5a, which is the last commit to touch that doc.

    New: relocation without locking writes to a file nothing reads

    This is the one thing I would treat as a defect rather than unfinished work, and it falls out of the design answer that LIBREFANG_CONFIG_PATH relocates and nothing more (crates/librefang-kernel/src/config.rs:658).

    Boot honours it: load_config(None) falls through to default_config_path() (config.rs:660), which returns $LIBREFANG_CONFIG_PATH when set. config_provenance resolves the same way, so GET /api/config/status reports source: /etc/librefang/config.toml.

    Nothing that writes does. Every persisting handler builds its own path as state.kernel.home_dir().join("config.toml") — config/manage.rs:1372, budget.rs:601, users.rs:1192, memory.rs:1711, providers.rs:1637, :1676, :2483, :2697, config/system.rs:264, channels.rs:1215 and :1315 — and Kernel::reload_config reads self.home_dir_boot.join("config.toml") (crates/librefang-kernel/src/kernel/config_reload_ops.rs:40).

    So with LIBREFANG_CONFIG_PATH=/etc/librefang/config.toml and the default mutable mode — exactly the Compose-bind-mount case the orthogonality argument was made to protect — every dashboard save lands in /data/config.toml, which the daemon never reads, while the status endpoint names a different file. The write returns 200 and is lost on restart.

    POST /api/config/reload then has two failure shapes: it 400s with Config file not found: /data/config.toml (config.rs:381), naming a path the operator never configured; or, if a pre-relocation config.toml is still sitting in the PVC — the normal case for anyone migrating into this feature — it silently swaps the live config to that stale file's contents. That is the criterion "invalid externally supplied configuration never partially replaces the last valid effective configuration", failing on the merely-relocated path.

    The comment at routes/channels.rs:1208-1213 shows the writers deliberately track kernel.home_dir() so they agree with reload_config(). That reasoning is sound; the problem is that both of them, not just one, are the odd ones out versus boot. The fix is to make reload_config and the write path resolve through default_config_path() (or a kernel-held config path captured at boot), not to move boot onto home_dir.

    Recommendation

    Keep this open. Concretely, the remaining items are: the dashboard provenance UI (section 4), the Kubernetes ConfigMap + Secret + checksum-rollout + rollback manifests (sections 1 and 6), the lock-or-classify decision on the channels / MCP / extensions / change-password writers, the include-file checksum semantics, and the missing integration cases.

    The config-path inconsistency above is worth its own issue rather than a checklist line here — it bites mutable mode with no managed mode in sight, so it should not wait on the rest of this RFC. I have not opened one; say the word and I will, or take it as an item here if you would rather keep it under the same umbrella.

  6. houko commented on Aug 24, 2026

    @houko
    Contributor

    Closing this as part of a backlog sweep, as not planned. @whatnick — roughly half of this RFC is on main and has been for months, but completed would be a lie about the other half, so this closes with the remainder written down rather than marked done.

    Re-verified against origin/main at 45e9bf06a.

    What landed

    #6717 (35d3ab59) — managed mode.

    • The two-env-var contract, kept orthogonal: LIBREFANG_CONFIG_PATH relocates, LIBREFANG_CONFIG_MODE=managed locks, each meaningful without the other. That is design question 1 answered as "both, independent".
    • guard_config_write() at crates/librefang-api/src/routes/mod.rs:361, returning exactly the shape §2 asked for — 423 Locked with {ok:false, error, code:"config_managed", source} (routes/mod.rs:376 and :396) — and reading the mode from the environment on every call, so a write cannot unlock its own file.
    • GET /api/config/status with mode / source / writable / checksum / modified_at.
    • Whole-config lock rather than field-level overlays (design question 2), and rollout-only updates (design question 4).
    • The managed-mode boot skips the migration write-back and logs one targeted warning, instead of retrying a doomed std::fs::write on every restart forever. That was not in the RFC; it fell out of the acceptance-criteria gap noted in the first review here.

    #6737 (5352cce5a) guarded the three provider handlers — set_provider_key, set_provider_url, set_default_provider — that #6717's own known-gaps list had missed, and corrected that list.

    #6982 (c6098ae2c) guarded PATCH /api/memory/config; #6988 (9797b3e0e) guarded POST /api/init.

    Eight guarded call sites in total, all firing before the handler touches the file, with integration coverage in crates/librefang-api/tests/config_managed_mode_test.rs.

    What remains

    1. The dashboard provenance surface (§4) has no consumer on main. Grepping crates/librefang-api/dashboard/src for config/status, ConfigStatus and config_managed returns nothing. So the criterion "the dashboard clearly marks managed settings and does not offer write controls for them" is unmet, and the UI still learns about managed mode by attempting a save and reading the refusal — the exact thing GET /api/config/status was added to avoid.

    This one has a home: open PR #7868, feat(config): surface managed mode in the dashboard and lock four remaining write domains — not a draft, not merged as of this comment. It carries the ConfigPage provenance banner and inert field rendering, and it also settles the lock-or-classify decision on the routes the guard never reached (change_password, the sidecar-channel configure/delete pair, the extension install/uninstall pair, and the MCP server routes under mcp_runtime_store = "file" only), with the deliberately-left-writable set written down alongside. If you want to follow one thread from this issue, follow that PR.

    2. The Kubernetes criteria were never started. deploy/kubernetes/ is README.md, base/kustomization.yaml, base/service.yaml, base/statefulset.yaml and secrets.example.yaml. There is no ConfigMap, and git grep for LIBREFANG_CONFIG_PATH or LIBREFANG_CONFIG_MODE under deploy/ returns nothing at all. So "a new installation can mount config.toml read-only outside /data and boot without an init-container copy" exists as prose in docs/operations/managed-config.md and nowhere else, and §6's requirement that the example place OAuth client secrets, API keys, the vault key and LIBREFANG_STATE_SECRET in Secrets has no artefact behind it. The checksum field a rollout annotation would key on already exists on the status endpoint; only the manifests, the Secret-reference example, the checksum-triggered rollout and the rollback procedure are missing.

    3. Declarative provisioning — the second half of this issue's title — does not exist. No provisioning/ tree, no resource reconciler, no per-resource provenance, no prune policy. The only provisioning paths in the repo are Grafana's under deploy/grafana/.

    4. Two design decisions are still open, and naming them is most of what is left here.

    Field-level ownership overlays versus whole-config. Currently whole-config, deliberately. The argument for keeping it that way: WRITABLE_EXACT_PATHS / WRITABLE_SECTION_PREFIXES is already a field-level axis and already load-bearing for security (external_auth. is absent from it because of the #3703 impersonation vector), so a second orthogonal ownership axis produces a matrix whose interesting cases are all corners. The argument against: a deployment that wants its IdP config immutable while leaving budgets editable has no way to say so today. Alternatives are (a) stay whole-config, (b) reuse the existing allowlist as the ownership axis rather than adding one, (c) a separate [managed] owns = [...] path list. Nobody has a deployment forcing the choice, which is why it is still open.

    Provenance for include files. The checksum is computed over the primary file's raw bytes, so editing an included file leaves it unchanged and an operator using the checksum to confirm a rollout landed gets a false negative. Alternatives: checksum the transitive closure of included files; refuse include entirely in managed mode; or document primary-file-only as intended and tell operators not to key rollouts on it. This has to be decided before anyone builds the checksum-annotation rollout in item 2, because it determines whether that annotation is trustworthy.

    5. One live defect that is not part of this RFC and should not have waited on it. Relocation without locking writes to a file nothing reads. Boot resolves through default_config_path() (crates/librefang-kernel/src/config.rs:660), which honours LIBREFANG_CONFIG_PATH. Every persisting handler instead builds state.kernel.home_dir().join("config.toml") — twenty sites, including routes/budget.rs:601, routes/users.rs:1192, routes/memory.rs:1775, routes/config/manage.rs:1375, routes/config/system.rs:145, routes/channels.rs:1330, routes/providers.rs:1767/:1806/:2638/:2844, routes/skills/extensions.rs:171/:241, routes/skills/mcp.rs:141/:166, server.rs:1172/:2281 — and Kernel::reload_config reads self.home_dir_boot.join("config.toml") (crates/librefang-kernel/src/kernel/config_reload_ops.rs:40).

    So with LIBREFANG_CONFIG_PATH set and the default mutable mode — the Compose bind-mount case the orthogonality decision was made to protect — every dashboard save returns 200, lands in /data/config.toml, is never read, and is lost on restart, while GET /api/config/status names a different file. POST /api/config/reload then either 400s naming a path the operator never configured, or silently swaps the live config to a stale pre-relocation file still sitting in the PVC. The fix is to route the writers and reload_config through default_config_path() (or a kernel-held path captured at boot), not to move boot onto home_dir.

    I said in an earlier comment that this deserved its own issue and did not file one. It still does — it bites mutable mode with no managed mode anywhere in sight, and it should not be buried in a closed umbrella. Anyone picking this up should file it standalone rather than reopening here.

    Closing note

    This is a backlog sweep, not a judgement on the RFC — the design questions in it were answerable from the code, which is unusual and is why the first half shipped cleanly. Items 2, 3 and 4 above are genuinely wanted; they are unbuilt, not unwelcome. Reopening this, or filing narrower issues for the Kubernetes manifests, the declarative-provisioning reconciler and the relocation defect, is welcome. PR #7868 remains the live thread for the dashboard half.

  7. reopened this on Aug 24, 2026
  8. houko commented on Aug 24, 2026

    @houko
    Contributor

    Reopened — closing this as not planned was a mistake on my part.

    not planned reads as "we do not intend to do this", which is not true of this issue.
    The work is either in flight in an open pull request or tracked as real remaining work, and the preceding comment says which.

    That comment's content stands: what landed, what remains, and where the work lives are all accurate.
    Only the closure was wrong, and the label it carried misrepresented the state to anyone reading the tracker.

    Issues covered by an open PR will close on their own when that PR merges, with the correct reason.

  9. houko commented on Aug 25, 2026

    @houko
    Contributor

    PR #7921 lands the second half of this RFC — the declarative provisioning the title names — plus the one open design question that was blocking the first half from being trustworthy. Deliberately Refs, not Closes; what remains is at the bottom.

    Re-verified as already shipped

    Three PRs merged after the last status comment here, and they close three of its five open items. Checked against origin/main at 3624319cc:

    So the accurate remaining set was: declarative provisioning (which did not exist in any form), and the two design decisions in item 4.

    What #7921 adds

    Declarative resource provisioning. LIBREFANG_PROVISIONING_PATH points at a deployment-owned tree of agents/*.toml, reconciled at boot into the registry. Stable identifiers (the manifest name, required to be present rather than defaulted), idempotent application, per-resource provenance persisted across restarts, drift reporting, kubectl apply-style adoption of an agent that already exists under a declared name, and an explicit prune policy whose default releases an orphan back to runtime ownership rather than deleting it — which is what makes removing a declaration reversible.

    Each declared agent is locked individually: nine routes answer 423 Locked with code: "resource_provisioned", in the same envelope shape config_managed uses. Operating the agent is untouched, and everything the tree does not declare stays fully mutable. GET /api/provisioning/status reports provenance, drift and every file the reconcile refused. This is the RFC's "provisioned resources should be locked individually while runtime-created resources remain mutable", built rather than sketched.

    The include checksum question, answered by fixing it. Of the three options the last comment listed, #7921 takes the first: the checksum covers the transitive closure. A deployment with no include keeps the exact digest it had, so #7902's shipped annotation and its kind assertion keep matching; with includes it becomes the digest of sha256sum output over the closure, reproducible from a shell. scripts/check-k8s-manifests.py consequently stops banning include — that ban existed precisely because the checksum was untrustworthy — and instead verifies that each included file is another key of the same ConfigMap.

    That matters here beyond tidiness: the last comment noted this had to be decided before anyone built the checksum-annotation rollout, because it determines whether the annotation is trustworthy. That rollout shipped in #7902 with the ban standing in for the answer. It is now answered properly.

    The field-level ownership question, answered rather than deferred. Ownership is expressed per resource, through provisioning, rather than per config field. That is the granularity a deployment has actually asked for — "the deployment owns these agents" — and it avoids layering a second orthogonal axis over WRITABLE_EXACT_PATHS, where the interesting cases would all be corners. config.toml stays whole-config. Recorded in docs/operations/managed-config.md so the decision is discoverable rather than folklore.

    What is left after #7921 merges

    Two things, both scoped and neither blocked on anything in this issue:

    1. channels/ and workflows/ provisioning. The RFC's tree names them; feat(provisioning): declarative resource provisioning and include-aware config checksums #7921 provisions agents/ only, and reports an unrecognised subdirectory as a failure rather than skipping it, so a deployment that tries provisioning/channels/ is told it is unsupported instead of believing it worked. They are out of scope for one PR for a substantive reason rather than size: neither is a matter of pointing the same scan at another subdirectory. Channels persist into config.toml itself, which the whole-config lock already covers, so provisioning them means deciding how a per-channel declaration composes with a file the deployment already owns. Workflows persist into a SQLite-backed registry carrying run state, so a reconcile has to define what happens to in-flight runs when a definition changes or is pruned — a question the agent path does not raise, because an agent's sessions survive a manifest swap.

    2. In-place reload of the provisioning tree. Boot-time only, matching the rollout-only contract §5 settled on for configuration. The same reasoning applies unchanged: a watcher has to handle Kubernetes' symlink-swap semantics and guarantee that a partially-written declaration never replaces a valid one, which is a second hard problem stacked on the first.

    Both deserve their own issues rather than keeping this one open, and I would rather file them than leave an umbrella that reads as unfinished when the thing it was opened for is built. Say the word and I will file them; if you would rather keep them here, this issue stays open and #7921 does not close it either way.

    The RFC itself — @whatnick's — is worth crediting again: the design questions in it were answerable from the code, which is why both halves shipped without a redesign.

  10. added
    has-prA pull request has been linked to this issue
    and removed
    has-prA pull request has been linked to this issue
    on Aug 27, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    awaiting-responseMaintainer replied — waiting for reporter/contributor feedbackneeds-triageAuto-applied when the issue title/body matched no area label — needs maintainer review

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions