Two JIT Gateway instances sharing leader election each enqueue billing audits in their own process. A follower that never becomes the BillingAuditService leader does not start its periodic flush timer. If it receives fewer than 100 events and then stops, its queue is inferred to be discarded without the service's shutdown drain.
This is a source-inspection finding, not an actual two-host HTTP reproduction. The evidence below is pinned to master commit 888e688d31729c60f1e37de0fb440f30db9f808d; the current feature branch retains this legacy JIT registration.
- Gateway billing registration, lines 51–55 registers a singleton
IBillingAuditService and wraps that same instance in leader election.
- BatchAuditServiceBase, lines 18–23 owns a
ConcurrentQueue<TEvent>; its constructor, lines 73–75 creates an initially inactive timer. StartAsync, lines 264–273 enables the 10-second flush timer.
- LeaderElectedServiceWrapper, lines 33–44 creates
_innerServiceCts and calls the inner StartAsync only on becoming leader. StopInnerServiceAsync, lines 66–94 calls the inner StopAsync only when that CTS exists.
- Wrapper disposal, lines 105–122 uses the same conditional stop, then disposes the inner service. BatchAuditServiceBase.Dispose, lines 332–354 disposes resources without draining the queue; the drain belongs to StopAsync, lines 285–308.
The full-batch path remains available on a follower: LogEventAsync/LogEvent, lines 131–173 triggers persistence at a queue depth of 100 independently of hosted-service startup. This finding concerns the remaining sub-threshold events and periodic flushing, not a claim that every follower audit is lost.
Reproduction to implement: start two JIT Gateways with shared Redis/PostgreSQL; confirm one stays the follower; send trackable requests only to that follower so its audit queue remains between 1 and 99; wait beyond the 10-second flush interval; gracefully stop the follower before it acquires leadership; independently query BillingAuditEvents using the known request IDs. The expected result is durable rows for every queued event, including after graceful shutdown.
Acceptance:
- Both instances persist their own sub-threshold audits during normal operation and graceful shutdown, including an instance that never becomes leader.
- Leadership changes do not duplicate or discard audit rows; the existing 100-event batching path remains covered.
- Database-failure requeue/recovery behavior remains covered.
- Shared retention work remains coordinated separately from flushing process-local queues.
Related closed tickets: #820 introduced the broader background-service coordination concern; #1016 concerns database flush failures and requeueing. Neither describes this follower lifecycle gap. Searches for "billing audit" follower, "BillingAuditService", and "audit" "leader" found no dedicated duplicate.
This is non-blocking for epic #1368's planned native path: its typed durable writer and per-instance hosted flush are intended to handle local queues independently, with unsupported retention disabled. That native work does not by itself fix or qualify the unchanged legacy JIT path.
Two JIT Gateway instances sharing leader election each enqueue billing audits in their own process. A follower that never becomes the BillingAuditService leader does not start its periodic flush timer. If it receives fewer than 100 events and then stops, its queue is inferred to be discarded without the service's shutdown drain.
This is a source-inspection finding, not an actual two-host HTTP reproduction. The evidence below is pinned to master commit
888e688d31729c60f1e37de0fb440f30db9f808d; the current feature branch retains this legacy JIT registration.IBillingAuditServiceand wraps that same instance in leader election.ConcurrentQueue<TEvent>; its constructor, lines 73–75 creates an initially inactive timer. StartAsync, lines 264–273 enables the 10-second flush timer._innerServiceCtsand calls the innerStartAsynconly on becoming leader. StopInnerServiceAsync, lines 66–94 calls the innerStopAsynconly when that CTS exists.The full-batch path remains available on a follower: LogEventAsync/LogEvent, lines 131–173 triggers persistence at a queue depth of 100 independently of hosted-service startup. This finding concerns the remaining sub-threshold events and periodic flushing, not a claim that every follower audit is lost.
Reproduction to implement: start two JIT Gateways with shared Redis/PostgreSQL; confirm one stays the follower; send trackable requests only to that follower so its audit queue remains between 1 and 99; wait beyond the 10-second flush interval; gracefully stop the follower before it acquires leadership; independently query
BillingAuditEventsusing the known request IDs. The expected result is durable rows for every queued event, including after graceful shutdown.Acceptance:
Related closed tickets: #820 introduced the broader background-service coordination concern; #1016 concerns database flush failures and requeueing. Neither describes this follower lifecycle gap. Searches for
"billing audit" follower,"BillingAuditService", and"audit" "leader"found no dedicated duplicate.This is non-blocking for epic #1368's planned native path: its typed durable writer and per-instance hosted flush are intended to handle local queues independently, with unsupported retention disabled. That native work does not by itself fix or qualify the unchanged legacy JIT path.