One gap in the claim query: attempt_count only moves when a live worker records a failure. A lease that expires because the worker died never increments it, and the claim UPDATE accepts expired LEASED rows without touching the count. So a payload that kills its worker, an oversized body that runs it out of memory or one that crashes the decrypt step, gets reclaimed every lease TTL indefinitely and never reaches max_attempts or the DLQ, taking a worker down each time. The fix fits in the statement you already have: count claims, not failures. Increment a claim_count in the same atomic UPDATE that bumps the fencing token, and route the row to the DLQ once it passes a limit. The audit table makes this case easy to spot, LEASED followed by LEASED with no ATTEMPTED in between, and the SQS comparison at the end already works this way, since maxReceiveCount counts receives rather than errors.