Repository navigation
RFC: add managed configuration and declarative provisioning mode #6695
Description
Activity
- addedneeds-triageAuto-applied when the issue title/body matched no area label — needs maintainer reviewAuto-applied when the issue title/body matched no area label — needs maintainer review
on Aug 1, 2026 This RFC describes current behaviour accurately, and the four design questions at the end are answerable from the code. Taking them in order, with the audit item first since it is also an acceptance criterion.
"Which existing write routes mutate deployment configuration versus legitimate runtime state?"
/api/config/setis not the only surface, and the gap is wider than the RFC assumes. Routes that persist intoconfig.tomltoday, each with a test that pins the behaviour:PUT /api/budget—budget_routes_test.rs:279is literally namedbudget_put_persists_to_config_toml.PUT /api/providers/{name}/budget—provider_budget_routes_test.rs:83, "persist seed config so PUT round-trips through a realconfig.toml".- User creation —
users_test.rs:524,users_create_refuses_to_overwrite_corrupt_config_toml. - Channel removal —
channel_remove_test.rs:3, "so theconfig.tomlrewrite lands in the sandbox". - Provider, skill, extension, and memory routes all reach the same file (
routes/{providers,skills/*,memory,backup,channels,users,config/*}.rs).
And one that is not a route at all, which is the sharpest problem for a read-only mount:
crates/librefang-kernel/src/config.rs:213writes the migrated config back to disk at boot whenevermigrated && file_version < CONFIG_VERSION. A ConfigMap mounted read-only makes thatstd::fs::writefail on the first boot after any config-schema bump. It currently degrades to atracing::warn!rather than aborting, so it will not crash — but it will re-run the migration on every subsequent boot, silently, forever. Managed mode has to make this an explicit, once-per-boot "migration required, config is deployment-managed" error rather than a warning that scrolls past.The RFC's instinct that "filesystem read-only errors must never be the enforcement mechanism" is right, and this is the concrete case that proves it.
"Should the public contract use
LIBREFANG_CONFIG_MODE=managed, infer fromLIBREFANG_CONFIG_PATH, or require both?"Both, kept orthogonal.
Relocating the file and locking it are different wants. An operator may reasonably want
config.tomloutsideLIBREFANG_HOMEwhile still editing it from the dashboard — a Compose deployment with a bind-mounted config directory, say. If the path implies the lock, that operator gets a read-only dashboard they never asked for, and the only way out is moving the file back.So:
LIBREFANG_CONFIG_PATHrelocates and nothing more.LIBREFANG_CONFIG_MODE=managedlocks, and is meaningful with or without a custom path. The RFC's own constraint — the mode must not be settable from inside the managed file — is satisfied either way, and keeping them independent means the Kubernetes manifest states its intent explicitly rather than encoding it in a path."Should managed mode lock the entire kernel configuration or allow field-level ownership overlays?"
Whole config, in the first implementation.
This codebase already has a field-level write model:
WRITABLE_EXACT_PATHSandWRITABLE_SECTION_PREFIXESincrates/librefang-api/src/routes/config/mod.rs:666. It is already load-bearing for security rather than convenience —external_auth.is deliberately absent from it because, per the comment atroutes/config/manage.rs:158, flipping an endpoint or the verification gate post-auth is the #3703 impersonation vector. There is also awritable ⊆ readableparity guard built on top of it.Layering a second, orthogonal ownership axis over that produces a two-dimensional matrix — writable-by-allowlist × owned-by-deployment — where the interesting cases are the corners, and where a future contributor adding a field has to reason about both. The security value of the existing allowlist comes from it being the single answer to "can this be written over HTTP". Keep it that way: in managed mode the answer is no, uniformly.
Field-level overlays are a reasonable v2 once someone has a deployment that actually needs one. Designing for it now costs clarity and buys nothing that has been asked for.
"Should the first release support rollout-only updates or also in-place reload?"
Rollout-only.
The reload classification the RFC wants to reuse already exists and is already documented:
build_reload_planincrates/librefang-kernel/src/config_reload.rsdecides hot-reload / restart-required / read-live, anddocs/operations/config-reload.mdis the ops-facing table, drift-guarded by a test. Exposing that classification through the API is cheap and worth doing in v1.Acting on it in place is not. In-place reload of a ConfigMap means handling the symlink-swap semantics of Kubernetes' atomic update — the RFC names this itself — plus deciding between
POST /api/config/reload,SIGHUP, a watcher, and a sidecar, and then guaranteeing that a partially-written or invalid file never replaces the last valid effective configuration. That is a second hard problem stacked on the first.A checksum annotation triggering a StatefulSet rollout works for restart-required fields too, which is the superset. Ship that, expose the classification, and let in-place reload be justified by a deployment that finds the rollout too slow.
One thing to add to the acceptance criteria
The criteria cover rejected writes but not the boot path. Worth adding: a managed-mode boot with a config file requiring schema migration produces a single actionable error and does not retry the write on every restart. That is the case the current
config.rs:213warning would otherwise hide.- addedawaiting-responseMaintainer replied — waiting for reporter/contributor feedbackMaintainer replied — waiting for reporter/contributor feedback
on Aug 4, 2026 🤖 PR #6717 just merged and appears to be the last open PR referencing this umbrella issue.
Dry-run flag from the umbrella-autoclose workflow. Close manually if correct, or reopen / add an unchecked task-list item if more work is planned. Delete this comment to let the bot re-flag later.
- addedauto-close-candidateBot thinks this is resolved; awaiting confirmationBot thinks this is resolved; awaiting confirmation
on Aug 7, 2026 #6717 merged as 35d3ab5, so here is what it covers against this RFC and what is left. It intentionally said
Refsrather thanCloses, and I think this issue should stay open — four acceptance criteria and one design question are still outstanding.Covered.
The two-env-var contract with relocation and locking kept orthogonal, which is design question 1 answered as "both, independent".
Server-side enforcement through a singleguard_config_write()(crates/librefang-api/src/routes/mod.rs:260) returning 423 Locked with the{ok:false, error, code:"config_managed", source}shape this RFC specified, refused before the handler reads the file — the reasoning for enforcing in-process rather than relying on a read-only mount is recorded atroutes/mod.rs:255.
GET /api/config/statusprovenance with mode, source, writable, checksum and modified_at.
Whole-config lock rather than field-level overlays (design question 2) and rollout-only updates (design question 4).
Plus one thing this RFC did not mention: a managed-mode load now skips the boot-path migration write-back and logs one targeted warning, instead of retrying a doomedstd::fs::writeon every boot forever.Still open.
-
The route audit is incomplete, and wider than the PR's own Known-gaps section says.
Three provider handlers still rewriteconfig.tomlin managed mode:set_provider_key(crates/librefang-api/src/routes/providers.rs:1452) throughpersist_default_modelat:1610and:1649,set_provider_url(:2408) throughupsert_provider_url(:2499) andupsert_provider_proxy_url(:2503), andset_default_provider(:2590) throughpersist_default_model(:2659), with the writes at:2769,:3016and:3061.
The Known-gaps list indocs/operations/managed-config.mdnames skills, memory andserver.rsinstead, and that list is wrong in both directions:routes/memory.rsandsrc/server.rscontain noatomic_writecalls at all,routes/skills/mod.rswrites skill secrets throughatomic_write_secret_file(:1052,:1073,:1085) rather thanconfig.toml, andproviders.rs— the one real gap — is not mentioned.
fix(api): lock the provider config routes in managed mode and correct the documented gap list #6737 guards the three handlers and corrects the list. -
RFC section 4 (the "Managed by deployment" badge and disabled write controls in the dashboard) is deferred.
-
RFC section 5 (in-place reload / ConfigMap watching) is deliberately out of scope in favour of rollout-only.
-
Field-level ownership overlays are deliberately out.
-
The Kubernetes acceptance criterion is met as prose only.
deploy/kubernetes/is untouched;docs/operations/managed-config.mdmentions ConfigMap and the checksum annotation in sentences at lines 7 and 74, with no manifest, no Secret-reference example, no OAuth/OIDC example and no rollback procedure.
Section 6's requirement that the example place OAuth client secrets, API keys, the vault key andLIBREFANG_STATE_SECRETin Secrets therefore has no artefact behind it yet. -
The design question about
includefiles is unanswered.
The checksum is computed over the primary file's raw bytes, so an edit to an included file leaves it unchanged and an operator using the checksum to confirm a rollout landed gets a false negative.
Integration coverage is also partial:
crates/librefang-api/tests/config_managed_mode_test.rshas three cases (status-mutable, status-managed, and config/set refused with the file left untouched), and there is nothing yet for the newly guarded budget and users routes, secret redaction, or a managed-mode reload failure.I will convert the six items above into a checklist in this issue body so
close-umbrella-on-last-pr.ymldoes not flag it, and drop the staleneeds-triage.-
- added a commit that references this issue
on Aug 9, 2026 - removedauto-close-candidateBot thinks this is resolved; awaiting confirmationBot thinks this is resolved; awaiting confirmation
on Aug 12, 2026 Re-audited this against
maintoday. The write guard from my 2026-08-08 comment has held up and grown, but it satisfies roughly the first half of the RFC, so this should stay open. Details, then the one new defect I found while checking.Verified as shipped
guard_config_write()lives atcrates/librefang-api/src/routes/mod.rs:361and returns exactly the shape section 2 asked for —423 Lockedwith{ok:false, error, code:"config_managed", source}— reading the mode from the environment on every call rather than caching it at boot.Eight call sites, all before the handler touches the file:
Handler Guard config_set(routes/config/manage.rs:1277):1386quick_init(routes/config/system.rs:144):190persist_budget(routes/budget.rs:592):596persist_users(routes/users.rs:1176):1184memory_config_patch(routes/memory.rs:1700):1704set_provider_key(routes/providers.rs):1481set_provider_url:2422set_default_provider:2630The last three are #6737 (
5352cce5a), which closed the gap I named last time.
Section 3 is shipped too:GET /api/config/status(routes/config/manage.rs:1143, registered atroutes/config/mod.rs:42) returnsConfigProvenancewithmode/source/writable/checksum/modified_at, and the managed-mode boot skips the migration write-back with one targeted warning (crates/librefang-kernel/src/config.rs:289).Not the whole RFC
Four acceptance criteria and the second half of the title are still unmet, and none of them is blocked on the guard.
Section 4, the dashboard, has no consumer at all. Grepping
crates/librefang-api/dashboard/srcforconfig/status,ConfigStatus,config_managedand423returns nothing. So the criterion "the dashboard clearly marks managed settings and does not offer write controls for them" is unmet, and the UI still learns about managed mode by attempting a save and reading the refusal back — precisely what the doc-comment onconfig_statussays the endpoint exists to avoid.The Kubernetes criteria have no artefact.
deploy/kubernetes/base/iskustomization.yaml,service.yaml,statefulset.yaml— no ConfigMap, and the statefulset's env block setsLIBREFANG_HOME=/datawith neitherLIBREFANG_CONFIG_PATHnorLIBREFANG_CONFIG_MODE. Nothing underdeploy/mentions either variable. So "a new installation can mountconfig.tomlread-only outside/dataand boot without an init-container copy" is documented prose only, and section 6's requirement that the example place the OAuth client secrets, API keys, vault key andLIBREFANG_STATE_SECRETin Secrets has nothing behind it.Declarative provisioning — the second half of this issue's title — does not exist. No
provisioning/tree, no resource reconciler, no per-resource provenance or prune policy. The onlyprovisioningpaths in the repo are Grafana's underdeploy/grafana/.The route audit is still incomplete by decision, not by oversight.
docs/operations/managed-config.md:117-158enumerates what still writesconfig.tomlin managed mode — the sidecar-channel configure/delete pair, the MCP server routes undermcp_runtime_store = "file", the extension install/uninstall routes unconditionally, andchange_passwordinserver.rs— and says locking them is a product decision. That decision is still outstanding, and it is the criterion "every API route that persists deployment configuration is identified and locked server-side".Integration coverage is now six cases in
crates/librefang-api/tests/config_managed_mode_test.rs(status mutable, status managed,config/set, memory patch, and the three provider routes). Still nothing for the guarded budget, users and quick-init paths, nothing for secret redaction, and nothing for a managed-mode reload failure.Two rows of that Known-gaps table are now stale in the other direction:
PATCH /api/memory/configandPOST /api/initare listed as gaps but are both guarded, since #6982 (c6098ae2c) and #6988 (9797b3e0e) respectively — both landed after5352cce5a, which is the last commit to touch that doc.New: relocation without locking writes to a file nothing reads
This is the one thing I would treat as a defect rather than unfinished work, and it falls out of the design answer that
LIBREFANG_CONFIG_PATHrelocates and nothing more (crates/librefang-kernel/src/config.rs:658).Boot honours it:
load_config(None)falls through todefault_config_path()(config.rs:660), which returns$LIBREFANG_CONFIG_PATHwhen set.config_provenanceresolves the same way, soGET /api/config/statusreportssource: /etc/librefang/config.toml.Nothing that writes does. Every persisting handler builds its own path as
state.kernel.home_dir().join("config.toml")—config/manage.rs:1372,budget.rs:601,users.rs:1192,memory.rs:1711,providers.rs:1637,:1676,:2483,:2697,config/system.rs:264,channels.rs:1215and:1315— andKernel::reload_configreadsself.home_dir_boot.join("config.toml")(crates/librefang-kernel/src/kernel/config_reload_ops.rs:40).So with
LIBREFANG_CONFIG_PATH=/etc/librefang/config.tomland the default mutable mode — exactly the Compose-bind-mount case the orthogonality argument was made to protect — every dashboard save lands in/data/config.toml, which the daemon never reads, while the status endpoint names a different file. The write returns 200 and is lost on restart.POST /api/config/reloadthen has two failure shapes: it 400s withConfig file not found: /data/config.toml(config.rs:381), naming a path the operator never configured; or, if a pre-relocationconfig.tomlis still sitting in the PVC — the normal case for anyone migrating into this feature — it silently swaps the live config to that stale file's contents. That is the criterion "invalid externally supplied configuration never partially replaces the last valid effective configuration", failing on the merely-relocated path.The comment at
routes/channels.rs:1208-1213shows the writers deliberately trackkernel.home_dir()so they agree withreload_config(). That reasoning is sound; the problem is that both of them, not just one, are the odd ones out versus boot. The fix is to makereload_configand the write path resolve throughdefault_config_path()(or a kernel-held config path captured at boot), not to move boot ontohome_dir.Recommendation
Keep this open. Concretely, the remaining items are: the dashboard provenance UI (section 4), the Kubernetes ConfigMap + Secret + checksum-rollout + rollback manifests (sections 1 and 6), the lock-or-classify decision on the channels / MCP / extensions / change-password writers, the
include-file checksum semantics, and the missing integration cases.The config-path inconsistency above is worth its own issue rather than a checklist line here — it bites mutable mode with no managed mode in sight, so it should not wait on the rest of this RFC. I have not opened one; say the word and I will, or take it as an item here if you would rather keep it under the same umbrella.
- added a commit that references this issue
on Aug 24, 2026 Closing this as part of a backlog sweep, as
not planned. @whatnick — roughly half of this RFC is onmainand has been for months, butcompletedwould be a lie about the other half, so this closes with the remainder written down rather than marked done.Re-verified against
origin/mainat45e9bf06a.What landed
#6717 (
35d3ab59) — managed mode.- The two-env-var contract, kept orthogonal:
LIBREFANG_CONFIG_PATHrelocates,LIBREFANG_CONFIG_MODE=managedlocks, each meaningful without the other. That is design question 1 answered as "both, independent". guard_config_write()atcrates/librefang-api/src/routes/mod.rs:361, returning exactly the shape §2 asked for —423 Lockedwith{ok:false, error, code:"config_managed", source}(routes/mod.rs:376and:396) — and reading the mode from the environment on every call, so a write cannot unlock its own file.GET /api/config/statuswithmode/source/writable/checksum/modified_at.- Whole-config lock rather than field-level overlays (design question 2), and rollout-only updates (design question 4).
- The managed-mode boot skips the migration write-back and logs one targeted warning, instead of retrying a doomed
std::fs::writeon every restart forever. That was not in the RFC; it fell out of the acceptance-criteria gap noted in the first review here.
#6737 (
5352cce5a) guarded the three provider handlers —set_provider_key,set_provider_url,set_default_provider— that #6717's own known-gaps list had missed, and corrected that list.#6982 (
c6098ae2c) guardedPATCH /api/memory/config; #6988 (9797b3e0e) guardedPOST /api/init.Eight guarded call sites in total, all firing before the handler touches the file, with integration coverage in
crates/librefang-api/tests/config_managed_mode_test.rs.What remains
1. The dashboard provenance surface (§4) has no consumer on
main. Greppingcrates/librefang-api/dashboard/srcforconfig/status,ConfigStatusandconfig_managedreturns nothing. So the criterion "the dashboard clearly marks managed settings and does not offer write controls for them" is unmet, and the UI still learns about managed mode by attempting a save and reading the refusal — the exact thingGET /api/config/statuswas added to avoid.This one has a home: open PR #7868,
feat(config): surface managed mode in the dashboard and lock four remaining write domains— not a draft, not merged as of this comment. It carries theConfigPageprovenance banner andinertfield rendering, and it also settles the lock-or-classify decision on the routes the guard never reached (change_password, the sidecar-channel configure/delete pair, the extension install/uninstall pair, and the MCP server routes undermcp_runtime_store = "file"only), with the deliberately-left-writable set written down alongside. If you want to follow one thread from this issue, follow that PR.2. The Kubernetes criteria were never started.
deploy/kubernetes/isREADME.md,base/kustomization.yaml,base/service.yaml,base/statefulset.yamlandsecrets.example.yaml. There is no ConfigMap, andgit grepforLIBREFANG_CONFIG_PATHorLIBREFANG_CONFIG_MODEunderdeploy/returns nothing at all. So "a new installation can mountconfig.tomlread-only outside/dataand boot without an init-container copy" exists as prose indocs/operations/managed-config.mdand nowhere else, and §6's requirement that the example place OAuth client secrets, API keys, the vault key andLIBREFANG_STATE_SECRETin Secrets has no artefact behind it. Thechecksumfield a rollout annotation would key on already exists on the status endpoint; only the manifests, the Secret-reference example, the checksum-triggered rollout and the rollback procedure are missing.3. Declarative provisioning — the second half of this issue's title — does not exist. No
provisioning/tree, no resource reconciler, no per-resource provenance, no prune policy. The onlyprovisioningpaths in the repo are Grafana's underdeploy/grafana/.4. Two design decisions are still open, and naming them is most of what is left here.
Field-level ownership overlays versus whole-config. Currently whole-config, deliberately. The argument for keeping it that way:
WRITABLE_EXACT_PATHS/WRITABLE_SECTION_PREFIXESis already a field-level axis and already load-bearing for security (external_auth.is absent from it because of the #3703 impersonation vector), so a second orthogonal ownership axis produces a matrix whose interesting cases are all corners. The argument against: a deployment that wants its IdP config immutable while leaving budgets editable has no way to say so today. Alternatives are (a) stay whole-config, (b) reuse the existing allowlist as the ownership axis rather than adding one, (c) a separate[managed] owns = [...]path list. Nobody has a deployment forcing the choice, which is why it is still open.Provenance for
includefiles. The checksum is computed over the primary file's raw bytes, so editing an included file leaves it unchanged and an operator using the checksum to confirm a rollout landed gets a false negative. Alternatives: checksum the transitive closure of included files; refuseincludeentirely in managed mode; or document primary-file-only as intended and tell operators not to key rollouts on it. This has to be decided before anyone builds the checksum-annotation rollout in item 2, because it determines whether that annotation is trustworthy.5. One live defect that is not part of this RFC and should not have waited on it. Relocation without locking writes to a file nothing reads. Boot resolves through
default_config_path()(crates/librefang-kernel/src/config.rs:660), which honoursLIBREFANG_CONFIG_PATH. Every persisting handler instead buildsstate.kernel.home_dir().join("config.toml")— twenty sites, includingroutes/budget.rs:601,routes/users.rs:1192,routes/memory.rs:1775,routes/config/manage.rs:1375,routes/config/system.rs:145,routes/channels.rs:1330,routes/providers.rs:1767/:1806/:2638/:2844,routes/skills/extensions.rs:171/:241,routes/skills/mcp.rs:141/:166,server.rs:1172/:2281— andKernel::reload_configreadsself.home_dir_boot.join("config.toml")(crates/librefang-kernel/src/kernel/config_reload_ops.rs:40).So with
LIBREFANG_CONFIG_PATHset and the default mutable mode — the Compose bind-mount case the orthogonality decision was made to protect — every dashboard save returns 200, lands in/data/config.toml, is never read, and is lost on restart, whileGET /api/config/statusnames a different file.POST /api/config/reloadthen either 400s naming a path the operator never configured, or silently swaps the live config to a stale pre-relocation file still sitting in the PVC. The fix is to route the writers andreload_configthroughdefault_config_path()(or a kernel-held path captured at boot), not to move boot ontohome_dir.I said in an earlier comment that this deserved its own issue and did not file one. It still does — it bites mutable mode with no managed mode anywhere in sight, and it should not be buried in a closed umbrella. Anyone picking this up should file it standalone rather than reopening here.
Closing note
This is a backlog sweep, not a judgement on the RFC — the design questions in it were answerable from the code, which is unusual and is why the first half shipped cleanly. Items 2, 3 and 4 above are genuinely wanted; they are unbuilt, not unwelcome. Reopening this, or filing narrower issues for the Kubernetes manifests, the declarative-provisioning reconciler and the relocation defect, is welcome. PR #7868 remains the live thread for the dashboard half.
- The two-env-var contract, kept orthogonal:
Reopened — closing this as
not plannedwas a mistake on my part.not plannedreads as "we do not intend to do this", which is not true of this issue.
The work is either in flight in an open pull request or tracked as real remaining work, and the preceding comment says which.That comment's content stands: what landed, what remains, and where the work lives are all accurate.
Only the closure was wrong, and the label it carried misrepresented the state to anyone reading the tracker.Issues covered by an open PR will close on their own when that PR merges, with the correct reason.
- added 6 commits that reference this issue
on Aug 24, 2026 PR #7921 lands the second half of this RFC — the declarative provisioning the title names — plus the one open design question that was blocking the first half from being trustworthy. Deliberately
Refs, notCloses; what remains is at the bottom.Re-verified as already shipped
Three PRs merged after the last status comment here, and they close three of its five open items. Checked against
origin/mainat3624319cc:- Item 1, the dashboard provenance surface (§4) — done, feat(config): surface managed mode in the dashboard and lock four remaining write domains #7868 (
a0acb2c9, merged 2026-08-24). It also settled the lock-or-classify decision the comment listed as outstanding, coveringchange_password, the sidecar-channel configure/delete pair, the extension install/uninstall pair, and the MCP server routes undermcp_runtime_store = "file". - Item 2, the Kubernetes criteria (§1 and §6) — done, feat(k8s): add a managed-config overlay and boot off the mount #7902 (
d520adbe, merged 2026-08-25).deploy/kubernetes/overlays/managed-config/now has the ConfigMap, the read-only mount outside/data, both env vars, thechecksum/configrollout annotation, an OAuth/OIDC secret-reference example and a rollback procedure — plus akindjob that boots the overlay for real and asserts the 423. - Item 5, relocation writing to a file nothing reads — done, fix(config): resolve config.toml once at boot and use that path everywhere #7886 (
e.../merged 2026-08-25) via a boot-resolvedconfig_path_bootthat every reader and writer shares.
So the accurate remaining set was: declarative provisioning (which did not exist in any form), and the two design decisions in item 4.
What #7921 adds
Declarative resource provisioning.
LIBREFANG_PROVISIONING_PATHpoints at a deployment-owned tree ofagents/*.toml, reconciled at boot into the registry. Stable identifiers (the manifestname, required to be present rather than defaulted), idempotent application, per-resource provenance persisted across restarts, drift reporting,kubectl apply-style adoption of an agent that already exists under a declared name, and an explicit prune policy whose default releases an orphan back to runtime ownership rather than deleting it — which is what makes removing a declaration reversible.Each declared agent is locked individually: nine routes answer
423 Lockedwithcode: "resource_provisioned", in the same envelope shapeconfig_manageduses. Operating the agent is untouched, and everything the tree does not declare stays fully mutable.GET /api/provisioning/statusreports provenance, drift and every file the reconcile refused. This is the RFC's "provisioned resources should be locked individually while runtime-created resources remain mutable", built rather than sketched.The
includechecksum question, answered by fixing it. Of the three options the last comment listed, #7921 takes the first: the checksum covers the transitive closure. A deployment with noincludekeeps the exact digest it had, so #7902's shipped annotation and itskindassertion keep matching; with includes it becomes the digest ofsha256sumoutput over the closure, reproducible from a shell.scripts/check-k8s-manifests.pyconsequently stops banninginclude— that ban existed precisely because the checksum was untrustworthy — and instead verifies that each included file is another key of the same ConfigMap.That matters here beyond tidiness: the last comment noted this had to be decided before anyone built the checksum-annotation rollout, because it determines whether the annotation is trustworthy. That rollout shipped in #7902 with the ban standing in for the answer. It is now answered properly.
The field-level ownership question, answered rather than deferred. Ownership is expressed per resource, through provisioning, rather than per config field. That is the granularity a deployment has actually asked for — "the deployment owns these agents" — and it avoids layering a second orthogonal axis over
WRITABLE_EXACT_PATHS, where the interesting cases would all be corners.config.tomlstays whole-config. Recorded indocs/operations/managed-config.mdso the decision is discoverable rather than folklore.What is left after #7921 merges
Two things, both scoped and neither blocked on anything in this issue:
-
channels/andworkflows/provisioning. The RFC's tree names them; feat(provisioning): declarative resource provisioning and include-aware config checksums #7921 provisionsagents/only, and reports an unrecognised subdirectory as a failure rather than skipping it, so a deployment that triesprovisioning/channels/is told it is unsupported instead of believing it worked. They are out of scope for one PR for a substantive reason rather than size: neither is a matter of pointing the same scan at another subdirectory. Channels persist intoconfig.tomlitself, which the whole-config lock already covers, so provisioning them means deciding how a per-channel declaration composes with a file the deployment already owns. Workflows persist into a SQLite-backed registry carrying run state, so a reconcile has to define what happens to in-flight runs when a definition changes or is pruned — a question the agent path does not raise, because an agent's sessions survive a manifest swap. -
In-place reload of the provisioning tree. Boot-time only, matching the rollout-only contract §5 settled on for configuration. The same reasoning applies unchanged: a watcher has to handle Kubernetes' symlink-swap semantics and guarantee that a partially-written declaration never replaces a valid one, which is a second hard problem stacked on the first.
Both deserve their own issues rather than keeping this one open, and I would rather file them than leave an umbrella that reads as unfinished when the thing it was opened for is built. Say the word and I will file them; if you would rather keep them here, this issue stays open and #7921 does not close it either way.
The RFC itself — @whatnick's — is worth crediting again: the design questions in it were answerable from the code, which is why both halves shipped without a redesign.
- Item 1, the dashboard provenance surface (§4) — done, feat(config): surface managed mode in the dashboard and lock four remaining write domains #7868 (
- addedhas-prA pull request has been linked to this issueA pull request has been linked to this issueand removedhas-prA pull request has been linked to this issueA pull request has been linked to this issue
on Aug 27, 2026
Summary
LibreFang currently treats
$LIBREFANG_HOME/config.tomlas writable application state.This works for local and interactive deployments, but it does not provide a clean infrastructure-as-code contract for Kubernetes operators who want configuration sourced from an immutable ConfigMap or another deployment-managed file.
Add a managed configuration mode that separates immutable operator-owned configuration from mutable runtime data, enforces the distinction in the API, and presents provisioned settings as read-only in the dashboard.
Motivation
The supported Kubernetes baseline stores both configuration and runtime state under
/data.An operator can copy a ConfigMap into the PVC during startup, but this creates ambiguous ownership: copying on every restart overwrites dashboard changes, while copying only on first boot allows configuration to drift away from the manifest.
Mounting the ConfigMap directly is also unsafe today because LibreFang expects
config.tomlto be writable and several API routes update it.Grafana-style provisioning provides a clearer model: deployment-managed resources are visible in the UI, marked with their provenance, and cannot be edited through interactive surfaces.
LibreFang should provide the same explicit contract rather than relying on filesystem permissions or UI-only disabling.
The immediate use case is OAuth/OIDC configuration under
[external_auth], where public provider metadata belongs in a ConfigMap and client secrets belong in Kubernetes Secrets.The design should apply to the complete kernel configuration rather than introducing a one-off OAuth path.
Current behavior
config.tomlis loaded fromLIBREFANG_HOMEand is expected to remain writable./api/config/setmutates the file and applies the config reload plan.[external_auth]is already deliberately excluded from the generic writable-path allowlist, but there is no global managed-mode status or provenance model./dataPVC.Proposed design
1. Separate configuration source from runtime state
Introduce an explicit configuration path independent of
LIBREFANG_HOME, for example:LIBREFANG_HOME=/datawould continue to own SQLite, agents, logs, credentials, and other mutable runtime state.The managed configuration file could then be mounted read-only from a ConfigMap without copying it into the PVC.
The mode must be selected outside the managed file itself so an API write cannot disable its own lock.
The default must remain the current mutable behavior for compatibility with desktop, local, Compose, and existing Kubernetes installations.
2. Enforce managed mode server-side
When managed mode is active, every endpoint that persists deployment configuration must reject the write consistently, preferably with
423 Lockedand a structured response such as:{ "error": "configuration is managed by the deployment", "code": "config_managed", "source": "/etc/librefang/config.toml" }The audit must include more than
/api/config/set.Password changes, quick initialization, provider setup, and any domain route that writes
config.tomlmust either be locked or explicitly classified as mutable runtime state.Filesystem read-only errors must never be the enforcement mechanism.
3. Expose configuration provenance
Expose authenticated metadata through the config API, either alongside the redacted config or through a dedicated status endpoint:
{ "mode": "managed", "source": "/etc/librefang/config.toml", "writable": false, "checksum": "sha256:...", "loaded_at": "..." }Do not expose secret values or sensitive filesystem details to unauthenticated callers.
The checksum should cover the effective non-secret configuration or have clearly documented semantics.
4. Make the dashboard accurately read-only
The dashboard should display managed configuration normally, add a visible “Managed by deployment” badge, disable or remove write controls, and explain where changes must be made.
The UI must consume server-provided capability/provenance metadata rather than infer managed mode from failed writes.
Operational actions that do not mutate deployment configuration should remain available.
5. Define update and reload behavior
Document one supported Kubernetes update workflow.
A ConfigMap hash annotation can trigger a StatefulSet rollout, which is the simplest initial contract and works for restart-required fields.
If in-place reload is supported, define whether it is triggered by
POST /api/config/reload,SIGHUP, filesystem watching, or a sidecar, and ensure updates are read atomically rather than observing ConfigMap symlink swaps mid-read.The API should report the same hot/restart/read-live classification already described by the reload planner.
A failed managed reload must retain the last valid effective configuration and expose an actionable error rather than partially applying the new file.
6. Keep secrets outside ConfigMaps
Managed configuration should continue to reference secret environment-variable names or secret-file paths.
The Kubernetes example must place OAuth client secrets, API keys, the vault key, and
LIBREFANG_STATE_SECRETin Secrets rather than serializing them intoconfig.toml.Example target deployment shape:
Declarative resource provisioning
The RFC should decide whether the first implementation covers only
KernelConfigor establishes a reusable provenance model for later Grafana-style resource provisioning, for example:A later resource reconciler would need stable identifiers, idempotent application, provenance, drift reporting, conflict handling, and an explicit deletion/prune policy.
Provisioned resources should be locked individually while runtime-created resources remain mutable.
This broader reconciler does not need to block the initial managed
config.tomlmode, but the initial API metadata should avoid making it impossible.Security considerations
Non-goals
Acceptance criteria
config.tomlread-only outside/dataand boot without an init-container copy.Design questions
LIBREFANG_CONFIG_MODE=managed, infer managed mode fromLIBREFANG_CONFIG_PATH, or require both?GET /api/config,GET /api/config/status, or a general capabilities endpoint?includefiles report provenance and checksums?Related work
runAsNonRootdeployments.config.tomlprovider model that motivates the first provisioning use case.