Skip to content

JIT Gateway follower can lose queued billing audits below the batch threshold #1581

Description

@nickna

Two JIT Gateway instances sharing leader election each enqueue billing audits in their own process. A follower that never becomes the BillingAuditService leader does not start its periodic flush timer. If it receives fewer than 100 events and then stops, its queue is inferred to be discarded without the service's shutdown drain.

This is a source-inspection finding, not an actual two-host HTTP reproduction. The evidence below is pinned to master commit 888e688d31729c60f1e37de0fb440f30db9f808d; the current feature branch retains this legacy JIT registration.

The full-batch path remains available on a follower: LogEventAsync/LogEvent, lines 131–173 triggers persistence at a queue depth of 100 independently of hosted-service startup. This finding concerns the remaining sub-threshold events and periodic flushing, not a claim that every follower audit is lost.

Reproduction to implement: start two JIT Gateways with shared Redis/PostgreSQL; confirm one stays the follower; send trackable requests only to that follower so its audit queue remains between 1 and 99; wait beyond the 10-second flush interval; gracefully stop the follower before it acquires leadership; independently query BillingAuditEvents using the known request IDs. The expected result is durable rows for every queued event, including after graceful shutdown.

Acceptance:

  • Both instances persist their own sub-threshold audits during normal operation and graceful shutdown, including an instance that never becomes leader.
  • Leadership changes do not duplicate or discard audit rows; the existing 100-event batching path remains covered.
  • Database-failure requeue/recovery behavior remains covered.
  • Shared retention work remains coordinated separately from flushing process-local queues.

Related closed tickets: #820 introduced the broader background-service coordination concern; #1016 concerns database flush failures and requeueing. Neither describes this follower lifecycle gap. Searches for "billing audit" follower, "BillingAuditService", and "audit" "leader" found no dedicated duplicate.

This is non-blocking for epic #1368's planned native path: its typed durable writer and per-instance hosted flush are intended to handle local queues independently, with unsupported retention disabled. That native work does not by itself fix or qualify the unchanged legacy JIT path.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions