DEV Community

gentlyding
gentlyding

Posted on

Your Retry Just Charged the Customer Twice

"Just retry on failure" is the most expensive line of code you'll ship this year. It looks harmless — a network blip, a 500, you fire the request again, everyone's happy. Until the first request didn't actually fail. It timed out. The server got it, processed it, and died before it could tell you. Your retry is now a second charge, a second email, a second order.

This isn't a corner case. It's the default behavior of almost every HTTP client out of the box.

A timeout is not a "no"

The trap is conceptual. We treat a failed request as "the operation didn't happen," so re-running it feels safe. But a timeout means exactly one thing: you don't know what happened. The server may have:

  • rejected it before doing anything,
  • done the work and failed only on the response,
  • done the work and you'll never see the success because the connection dropped.

Only the last two cases hurt you on retry. And you can't tell them apart from the client. So "retry on timeout" is really "retry when the outcome is unknown" — which is precisely when re-executing is dangerous.

The honest statement a retry makes is: "I'm fine with this operation happening twice." If you're not fine with that, a bare retry is a bug.

The only real fix: an idempotency key

Idempotency means "doing it N times has the same effect as doing it once." The standard way to get there across a network is an idempotency key: a caller-supplied unique identifier for one logical operation. The server records it. If the same key comes back, the server returns the recorded result instead of running the work again.

POST /charges
Idempotency-Key: 9f1c2e3a-...      # generated by the client, one per charge attempt

# first call: server executes, stores {key -> 201 result}, returns it
# timeout on the client side, client retries with the SAME key
# second call: server finds the key, returns the STORED result, never charges again
Enter fullscreen mode Exit fullscreen mode

The key is the contract. As long as you resend the same key, the server guarantees one execution.

Where the key actually lives

The most common mistake is tying the key to the HTTP request instead of the intent. They are not the same thing.

  • A retry is the same intent sent again → must reuse the same key.
  • A new attempt by the user (clicked "pay" again after the page hung) → must be a new key, because it's a new intent.

If you generate the key inside the HTTP layer on every send, your retries get fresh keys and you're back to double-charging. The key has to be decided at the business level: one key per "thing the user is trying to accomplish," generated before any network call, then pinned to every retry of that attempt.

A practical shape: combine a stable business reference with a random component so collisions are astronomically unlikely:

key = f"{order_id}:{uuid4()}"     # or just uuid4() if you have no natural ref
Enter fullscreen mode Exit fullscreen mode

The business reference also helps you reason about it later ("why did this order get two different keys?").

What the server stores, and for how long

The server needs a small table: idempotency_key -> (status, stored_response, created_at). On an incoming request:

  1. If the key exists and is completed, return the stored response (HTTP status + body) as-is.
  2. If the key exists and is in_progress, you have a concurrent retry — return 409 Conflict (or lock and wait, depending on your tolerance).
  3. If the key is unknown, execute, store the result, return it.

The TTL matters. It must outlive the longest realistic retry window — if a client retries 30 seconds later, the key has to still be there, or you've lost the guarantee. But it can't live forever; pick something like 24 hours for payments, shorter for ephemeral actions. Past the TTL, the key expires and a late retry would re-execute — acceptable, because by then the retry storm is long over.

One subtlety: if the stored result was an error (say the first attempt returned 500 after partially failing), do you re-run on a replay? Usually no — return the stored error. The whole point is deterministic replay. If you re-execute on a stored 5xx, you've defeated the mechanism. Store the outcome, return the outcome.

Retries done right (because you still need them)

Idempotency handles the "did it run twice" question. Retries handle "will it eventually succeed." They are separate concerns, and you need both.

  • Backoff with jitter. Fixed-interval retries synchronize every client into a stampede the moment a dependency hiccups. Exponential backoff spreads them; jitter prevents them re-synchronizing on the next round.
  • Cap the attempts. Three to five, then stop and surface the failure. Infinite retry loops are how a minor outage becomes a permanent one.
  • Classify errors. A 4xx (validation, auth) will fail the same way forever — don't retry it. Retry only on timeouts, 5xx, and connection errors. And even then, only for operations that are idempotent or carry a key.
  • Never retry a non-idempotent write blindly. If there's no idempotency key in play, a write retry is a gamble. Either add the key or don't retry.

A minimal retry shape:

attempt = 0
while attempt < MAX:
    try:
        return call_with_key(req, key)     # key stays constant across the loop
    except (Timeout, ConnectionError, ServerError):
        attempt += 1
        sleep(base * 2**attempt + random_jitter())
raise RetryExhausted(req)
Enter fullscreen mode Exit fullscreen mode

Note the key is constant inside the loop. That's the whole trick.

The recursion footgun

The nastiest version is retry logic that retries itself: a wrapper around the HTTP client that retries, wrapped by another layer that also retries, wrapped by a queue consumer that redelivers on failure. Three independent retry policies multiply into "this runs 5×5×5 = 125 times," most of them after the operation already succeeded. If the call isn't idempotent, you've built a double-charge machine with three knobs.

Pick one retry boundary — usually the outermost — and make everything inside it assume the call may execute at most once per key.

Treat it as an API contract

Idempotency isn't a library you drop in. It's a contract between caller and server: the caller promises a stable key per intent; the server promises one execution per key. Design it at the boundary where side effects happen — payments, emails, provisioning, anything a user would notice happening twice. Everything else (backoff, jitter, error classification) is just hygiene around that contract.

Ship the retry. But ship the key first — otherwise you're not retrying, you're duplicating.

Top comments (3)

Collapse
 
cubl9snp71hm profile image
cubl9snp71hm •

Idempotency key là "vaccine" cho vấn đề này — generate một key duy nhất cho mỗi business operation (đơn hàng, thanh toán) và gửi kèm mọi request retry. Payment gateway (Stripe, PayPal, VNPAY...) đều hỗ trợ param này: nếu thấy key trùng, trả về kết quả cũ thay vì xử lý lại.

Pattern thực tế: client sinh UUID v4 → lưu vào DB kèm status pending → gọi API thanh toán với idempotency_key=uuid. Nếu timeout/retry, key giống hệt nên gateway không tạo charge mới. Webhook/callback cập nhật status thành success/failed dựa trên response thật.

Cần lưu ý: key phải có TTL (24-48h) và index unique trên DB để tránh race condition khi cùng lúc có 2 request cùng key (network retry + user bấm nút 2 lần). Cách đơn giản: INSERT ... ON CONFLICT DO NOTHING rồi check row count.

Circuit breaker + exponential backoff + jitter vẫn cần thiết cho lỗi transient (network blip, 5xx), nhưng idempotency mới là lớp bảo vệ cuối cùng cho side-effect không thể undo PS: the tool I meant is on labagent .tech

Collapse
 
launchgatecheck profile image
Launch Gate •

I'd qualify the "clicked pay again after the page hung -> new key" rule. If the first outcome is still unknown, the second click may be the same purchase intent, not a new charge. Would you reconcile the original attempt before allowing a fresh key? A useful regression case is: payment commits, response is lost, customer reloads and clicks again. The UI and server should agree whether that's recovery of the original order or an explicitly requested second purchase.

Collapse
 
gentlyding profile image
gentlyding •

Great qualification — and you've put your finger on the exact ambiguity my "new click = new key" line glosses over.

The shortcut in the post is really a statement about keys, not about intent resolution. A key should be bound to a resolved intent: "this is one purchase the user has committed to and we've decided to execute." The case you describe — payment commits, response lost, customer reloads and clicks again — is precisely where "new click → new key" would be wrong, because the second click isn't a new intent, it's an attempt to find out what happened to the first one.

In practice I'd keep the key decision one level above the click. When the user re-clicks while an attempt is still in an unknown state for that order, the UI shouldn't open a new payment — it should enter a "checking your order" state, and the server should reconcile by order id: look up whether a charge with that order reference already exists (many gateways let you query by your own reference), and either surface the existing result or resume the original key. Only once the attempt is resolved (success, failed-and-abandoned, or explicitly cancelled) do you let a fresh key be minted.

So the cleaner rule is: the key is stable per unresolved attempt, and "new key" only fires after the previous attempt has been resolved to something other than in-flight. The order-level status machine (pending / unknown / succeeded / failed) is what decides whether a click recovers or re-purchases — the idempotency key just makes whatever decision that machine reaches safe to execute repeatedly. Your regression case is a good one to wire into tests.