快取與配額設計案例研究
一個在嚴格呼叫預算下,把對推論 API 的實際請求壓到最低的應用層快取閘道。
每次打 upstream 都會消耗昂貴的推論算力。因此我把目標設得更嚴:不只是「維持在 1,000/天以下」,而是將真實呼叫次數降到最低。系統必須用 1,000/天的預算服務每日 1,000,000+ 筆請求。
36 種參數組合(4 個時段 × 3 家飯店 × 3 種房型)× 每天 288 個 5 分鐘窗口:
遠超過 1,000 次的預算上限。我們需要 batch fetching + request coalescing + 硬性配額上限。
完整架構:client → Rails app(4 層)→ Redis + upstream。每一層在下方各自的區塊中詳細說明。
這個實作是用「一台 Redis」來搶鎖。這裡先誠實講一個極限:少數情況下,可能同時出現兩個 leader(兩個請求都跑去打 upstream)。原因有兩種,分開看:
原因一:leader 卡住了(一台 Redis 就會發生)
搶到鎖的請求,鎖設 7 秒自動過期(免得它掛掉、鎖卡死)。但如果它做到一半卡住超過 7 秒(比如 Ruby 的 GC 暫停、或 upstream 太慢),鎖就先過期了 → 下一個請求以為沒人在做,搶到鎖 → 變成第二個 leader。
原因二:Redis 換手(只有多台時才會)
如果之後改用 Sentinel/Cluster 那種多台備援,主機掛掉、切換到備機的瞬間,鎖可能還沒同步過去 → 備機上沒這把鎖 → 又冒出第二個 leader。
這兩種各自誰能解?
原因二 → 用 etcd/Consul 這種「共識演算法」的協調器,能保證換手不掉鎖。
原因一 → 共識演算法也救不了:卡住的請求醒來,一樣以為自己還握著鎖。而且要注意 —— 就算加了 fencing token,它擋的也只是「舊 leader 的寫入污染資源」,擋不了那第二次 upstream 已經打出去(要「不打第二次」得靠冪等,且上游要支援)。
我的選擇:不追求「絕對只有一個 leader」。
因為 upstream 多打一次其實沒差(拿到的是同一份資料)。真正的底線在下一層 —— QuotaGuard 用原子計數硬性鎖死每天 1000 次上限,就算偶爾兩個 leader,每日總量也絕不會破。所以這裡的鎖是拿來「省成本」的,不是拿來「保證正確」的;正確性交給 QuotaGuard。
補充:Redlock(用多台 Redis 搶過半數)比單一 Redis 的容錯好一點,但在時鐘偏移、GC 暫停下仍有已知問題,詳見 Martin Kleppmann 的分散式鎖分析。
app/services/rate_cache.rb
Upstream API 成本昂貴(每天 1,000 次上限)。相同的 (period, hotel, room) 組合在一段時間內會回傳相同的房價。
沒有 cache,每個請求都要打 upstream。每天 1,000,000 筆請求,配額會很快燒光。
以 (period, hotel, room) 為 key,設定 5 分鐘 TTL 存入 cache。在此時間窗口內,相同組合的請求全部從 cache 回傳 — 不打任何 upstream。
cache key: "pricing:Summer:FloatingPointResort:SingletonRoom" TTL: 5 minutes Total keys: 36 (4 periods × 3 hotels × 3 rooms)
app/services/lock_and_fetch_rate.rb
Cache 每 5 分鐘過期一次。過期瞬間,可能有大量請求同時進來,全部遇到 cache miss。
沒有協調機制,每個 concurrent cache miss 請求都會各自打 upstream。50 個並發請求會消耗 50 個配額單位,而不是 1 個。
使用 Redis lock(SET NX)搭配全域 lock key(lock_and_fetch_rate:pricing:bulk_fetch)選出一個「leader」。Leader 打 upstream 並填滿全部 36 個 cache 項目;其餘「followers」等待並從剛更新的 cache 讀取。
三個值形成一條計時鏈 — 每個值都依賴前一個:
沒有 timeout,upstream 如果卡住,請求會永遠阻塞並佔用 Puma worker。5 秒讓 upstream 有足夠時間正常回應,但在 upstream 掛掉時能快速釋放 worker。
Leader 可能在 4.9s 收到 upstream 回應並寫入 cache,但 follower 在 5.0s 就放棄了 — 白等一場。多 1 秒的緩衝處理這個邊界情況。
正常情況下,ensure 會在 leader 完成後立即刪除 lock。但若程序強制崩潰(如 kill -9),ensure 不會執行,lock 就卡住了。7 秒自動過期是最後的安全網。
必須比 FOLLOWER_TIMEOUT(6s)更長。若更短,lock 在 followers 還在等待時就過期 → 新請求搶到 lock → 成為第二個 leader → 再次打 upstream → 浪費配額。
1ms 太快,會大量打 Redis。1s 太慢,leader 完成後 follower 最多還要等 1 秒。50ms 是好的平衡點,最多增加 50ms 的額外延遲。
app/services/quota_guard.rb
搭配 Batch Fetch,理論最大值只有 288 次/天 — 遠低於 1,000 的預算。正常運作下,QuotaGuard 永遠不會觸發。它是一個本地保險絲,用來防範不該發生但可能發生的情況 — 例如讓每次 cache miss 都變成 upstream 呼叫的 bug:
rate 改名成 price。解析失敗、cache 一直空著,每個請求都打 upstream,直到 upstream 自己開始拒絕。RateCache.write_all 的 key 名稱打錯字。寫入悄悄成功但讀取永遠對不上 → 同樣的失控迴圈。正常情況下第 1,001 次呼叫會在前端就被 429 擋掉 — upstream 根本不會為它跑那段昂貴的推論。但萬一這個限制器設定出錯(bug、重新部署、staging 的怪狀況),溢出的請求就會真的打到計算成本極高的推論、燒掉真實算力。QuotaGuard 讓我們不論 upstream 的強制狀態如何都守住 1,000 次/天的合約。
接受的 trade-off:
DAILY_LIMIT 也得跟著改。多個 Puma worker 同時處理請求。簡單的「讀取 → 檢查 → 遞增」存在競態條件:兩個 worker 可能同時讀到 999,同時通過檢查,然後各自打 upstream。
// Race condition example: Worker A: read count → 999 → passes check → calls upstream // count is now 1000 Worker B: read count → 999 → passes check → calls upstream // EXCEEDED!
Redis INCR 是原子操作 — 在單一步驟內遞增並回傳新值。每個請求取得唯一的序號,恰好 1,000 次呼叫通過,沒有競態條件。
daily key: "quota:2026-04-22" // auto-expires at midnight Redis INCR → 1000 → pass ✓ Redis INCR → 1001 → raise ExhaustedError ✗
INCR 是原子操作,不可能超出。降低緩衝只在有逐漸減速策略時有意義。我們用硬性停止加上 Retry-After 到午夜,所以設更低的上限只是浪費配額。
Upstream rate API 文件沒有說明每日配額何時重置。本實作假設UTC 午夜 — Redis key 使用 quota:YYYY-MM-DD(UTC),Retry-After header 計算到 UTC 00:00 的秒數。若實際重置時間不同,只需修改 key 的日期邏輯與 Retry-After 計算。
lib/rate_api_client.rb → get_all_rates
Upstream API 支援在單一請求中查詢多個房價。總共有 36 種組合(4 個時段 × 3 家飯店 × 3 種房型)。
沒有 batch,每次 cache miss 只取一個房價 — 每個房價消耗一個配額單位。即使有 LockAndFetchRate,最糟情況仍是 36 次 cache miss × 288 個窗口 = 10,368 次/天。
任何 cache miss 發生時,透過 RateApiClient.get_all_rates 以單一 API 呼叫取得全部 36 個房價,再由 RateCache.write_all 一次填滿所有 cache key。一個配額單位讓整個 cache 在接下來 5 分鐘內保持有效。
Cache miss for "Summer:FloatingPointResort:SingletonRoom" → fetch 1 rate → write 1 cache entry → 1 quota unit for 1 rate
最糟情況:10,368 次/天
Cache miss for "Summer:FloatingPointResort:SingletonRoom" → fetch 36 rates (all combos) → write 36 cache entries → 1 quota unit for 36 rates
最糟情況:288 次/天(= 理論最小值)
每項決策列出考量的選項、最終選擇,以及有意識地接受的 trade-off。點擊展開。
Batch fetch 本身不是選擇題 — 數學就是這樣強迫的:36 種組合 × 每天 288 個 cache 窗口 = 10,368 次呼叫,遠超 1,000 的預算。所以每次 upstream 呼叫必須一次回傳全部 36 個房價。在 5 分鐘 TTL 下,288/天是上限的數學下限 — 沒有任何符合規範的設計在尖峰時能做得更少。真正值得思考的是何時觸發這個呼叫:第一次 cache miss 時(lazy — 下限為 0),還是按排程(prewarming — 固定 ≥ 288)。
| 選項 | 行為 |
|---|---|
| 定期 prewarming 例如 Sidekiq cron 每 4 分鐘執行 |
Cache 始終保持熱,每個請求都即時回應。但:即使零流量也消耗 360 次/天;需要額外基礎設施(Sidekiq + scheduler)。 |
| Lazy on-demand ✓ | Upstream 成本與實際需求成比例 — 沒人查詢時 0 次,滿載時最多 288/天。不需要額外基礎設施。 |
為什麼 lazy 適合這個工作負載:
Prewarming 是正確選擇的情況:流量可預測且穩定、延遲 SLA 嚴格(任何 cache miss 都不可接受),或配額預算充裕。在那些工作負載下,prewarming 勝出。不同限制條件,不同答案。
接受的 trade-off:第一個遇到過期 cache 的使用者要承擔小額的延遲成本(等待 upstream 呼叫)。之後在同一個 5 分鐘窗口內的所有請求都是即時 cache 命中。
| 選項 | 被排除的原因 |
|---|---|
MemoryStore | 每個 Puma worker 各自的記憶體 — 命中率除以 worker 數量,無法滿足 10:1 的需求。 |
SolidCache | 基於 SQLite;沒有原子 INCR 可用於配額;row-level lock 爭用。 |
| Redis ✓ | 跨 worker 共享;原子操作(SET NX、INCR、EXPIREAT)用一個依賴項解決 cache、lock 和配額三個問題。 |
Trade-off:docker-compose.yml 多一個服務。值得 — Redis 在 Rails production 環境中極為普遍。
沒有合併機制,50 個請求的爆發會在毫秒內燒掉 50 個每日配額單位。Redis SET NX 在所有 Puma worker 中選出一個 leader;其餘 followers 輪詢 cache。
// LockAndFetchRate pseudocode: token = SecureRandom.uuid // per-leader owner id (safe release, NOT a fencing token) if SET NX "lock:bulk_fetch" = token (TTL 7s) → leader yield // fetch upstream, write cache Lua: if GET == token then DEL // release only if still ours else → follower: poll cache every 50ms up to 6s
Owner id(安全釋放):每個 leader 持有一個唯一的 UUID。釋放時,Lua script 先確認「還是我的?」才刪除 — 否則一個緩慢的 leader 可能誤刪新 leader 的 lock,讓第三個請求趁虛而入。注意:這是 owner id,不是 fencing token — 它只防「刪錯別人的鎖」,擋不住過期後醒來的舊 leader 重複打 upstream(那要靠 fencing token + 資源檢查,或冪等)。
單一 Redis 假設:兩個 leader 可能短暫並存,成因有二:(1)leader 卡住(GC 暫停/upstream 慢過 lock TTL)→ lock 過期 → 第二個 leader,單一 Redis 就會發生;(2)Sentinel / Cluster 下 primary failover 遺失 lock 寫入。基於 Raft 的協調器(etcd / Consul)只解決成因 2,解決不了成因 1 — 那唯一的解是被保護資源檢查的 fencing token。這裡選擇不做:lock 是效率機制,正確性交給 QuotaGuard 的原子計數(fail-closed)— 即使兩個 leader 短暫並存,每日上限仍然成立。
Upstream 的限制是每天 1,000 次 — 不是每秒。所以我們需要一個在午夜重置的簡單計數器。
Batch Fetch 已將理論最大值壓在 288 次/天 — 遠低於預算。正常運作下,QuotaGuard 永遠不會觸發。
真正的殘餘價值是防禦性的客戶端行為:即使 upstream 自身的 rate limiter 因 bug、重新部署或 staging 問題而失效,我們仍將自己的呼叫速率限制在 1,000/天。在那種邊界情況下,溢出的請求會打到 upstream 那台計算成本極高的推論引擎並燒掉真實算力。QuotaGuard 防止了這件事。
接受的 trade-off:
DAILY_LIMIT 也需要更新。計數器必須在所有 Puma worker(獨立 process)間共享,而 INCR 是原子操作 — 它在單一步驟內遞增並回傳新值,所以兩個 worker 不可能同時讀到 999 並同時通過檢查。
quota:YYYY-MM-DD 執行 INCR — 原子操作,跨 workerINCR 是原子操作 — 不可能超出refund_quota! 將 DECR 包在 Lua script 中,只在計數器大於零時執行。沒有這個,在任何 INCR 之前爆發的 401 會把計數器推成負數,讓後續的 INCR 悄悄超出 1,000。配額在 UTC 午夜重置。Upstream 文件沒有說明每日限制何時重置,所以我們假設 UTC 00:00。Redis key 使用 quota:YYYY-MM-DD(UTC),Retry-After 計算到 UTC 午夜的秒數。如果實際重置時間不同,只需更新這兩個值。
當依賴項目掛掉(Redis、upstream、lock),服務應該怎麼辦?有四種常見策略 — 每種優先順序不同。
| 策略 | 行為 | 保護 | 犧牲 |
|---|---|---|---|
| ✓ Fail-closed 我的選擇 |
依賴掛掉 → 回傳 5xx,永不打 upstream | Upstream 合約(1,000/天上限) | 自身服務可用性 |
| Fail-open | 依賴掛掉 → 繞過保護,直接打 upstream | 自身可用性 | 配額快速耗盡;thundering herd 無法阻擋 |
| Graceful degradation | 回傳過期 cache、預設值或部分回應 | 使用者體驗 | 資料即時性 / 正確性 |
| Fail-fast + retry | 回傳 5xx,預期客戶端以 backoff 重試 | 與 fail-closed 相同,但將重試責任推給呼叫方 | 假設客戶端行為良好 |
為什麼這個服務選 fail-closed:兩條規範要求強迫了這個選擇。
排除這兩個後,fail-closed 是唯一誠實的選擇。Fail-fast + retry 功能上等效,但將正確性責任轉嫁給客戶端;對 proxy 來說,自己承擔這個責任更乾淨。
我會選擇不同策略的情況:
正確的選擇永遠取決於「錯誤資料」還是「沒有資料」傷害更大。對定價來說,錯誤資料的代價更高。
所有 upstream 失敗 — 401、5xx、timeout — 都回傳 400 Bad Request。客戶端無法區分請求本身有問題還是伺服器端失敗。
每種失敗對應不同的 HTTP 狀態碼。客戶端清楚知道發生了什麼,以及該修正請求、重試還是升級處理。
{"message":"Failed to process rates...","status":"error"}rate 欄位來自 1,001 次真實測試呼叫 — 完整模式表格見 UPSTREAM_BEHAVIOR.md。
Anti-Corruption Layer (ACL) RateApiClient 吸收所有 upstream 的不穩定行為,對外暴露乾淨的型別化介面。PricingService 永遠不會直接接觸原始 HTTP 回應:
Integer| 情境 | 誰拋出 | 客戶端狀態碼 | 客戶端訊息 | 客戶端動作 |
|---|---|---|---|---|
| 無效參數 | Controller | 400 | Missing required parameters / Invalid period | 修正請求 |
| 每日配額耗盡 | QuotaGuard | 429 | Daily quota reached — requests resume at midnight UTC | 等待 Retry-After |
| Upstream 回傳 HTTP 429(token 限制) | RateApiClient | 429 | Upstream API quota exhausted — token limit reached | 等待 Retry-After |
| 無效 token(401) | RateApiClient | 502 | Upstream authentication failed | — |
| Upstream 200 但 body 為錯誤 | RateApiClient | 502 | Upstream error | — |
| Upstream 200 但缺少 rate 欄位 | RateApiClient | 502 | Upstream error | — |
| Upstream 5xx | RateApiClient | 502 | Upstream error | — |
| Upstream 逾時 | RateApiClient | 504 | Upstream timed out | — |
| Follower 等待逾時 | LockAndFetchRate | 504 | Request timed out waiting for upstream result | — |
| 非預期錯誤 | StandardError catch-all | 500 | 非預期錯誤 | — |
兩種 429 回應都包含 Retry-After header(到 UTC 午夜的秒數)。其他所有錯誤回應只回傳狀態碼和訊息 — 不提供重試指引。
每個請求輸出一行 JSON log,包含:method、path、status、duration、cache_status("hit" 或 "miss")和 upstream_latency_ms。
// Example log line: {"method":"GET","path":"/api/v1/pricing","status":200,"duration":12.3, "cache_status":"hit","upstream_latency_ms":null}
為什麼選 Lograge 而不是付費服務(Datadog、New Relic):Lograge 免費、零設定,隨 Rails 提供。對這樣的案例研究,結構化 JSON log 足以展示可觀測性的思維。在 production 環境中,這些 JSON 行可以直接導入任何 log 聚合器。
為什麼重要:「你怎麼知道這個服務在 production 是健康的?」有了具體答案 — 篩選 cache_status=miss,觀察 upstream_latency_ms。
所有測試檔案遵循一致的段落結構:Happy path → Fail path → Edge/Boundary → Isolation。用 SimpleCov 量測覆蓋率 — 目前 100%(排除 scaffold 檔案)。
為什麼選 WebMock 而不是 .stub:直接 stub RateApiClient.get_all_rates 會繞過 HTTP 層,隱藏 timeout、header 和 retry 行為。WebMock 在 socket 層攔截,產生更真實的測試。
每個依賴項目失敗時會發生什麼,以及為什麼每日 1,000 次上限仍然成立。
| 什麼失敗了 | 內部發生什麼 | 客戶端看到 | 對 Upstream 的影響 | 恢復方式 |
|---|---|---|---|---|
| Redis 完全停機 例如崩潰、OOM、網路分區 |
Rails 內建的 cache 錯誤處理器吞掉 Redis 錯誤並回傳 nil → Rails.cache.write(..., unless_exist: true) 回傳 nil → LockAndFetchRate 進入 follower 分支 → poll_for_result 看到 nil 持續 6s → 拋出 LockAndFetchRate::TimeoutError |
504 "Request timed out waiting for upstream result" | 0 次 — 無人取得 lock,upstream 永遠不會被觸及 | Redis 恢復後服務立即回復(已由壓力與混沌測試中的 Redis 崩潰時 fail-closed 測試驗證) |
| Redis 重啟 例如部署、OOM、容器重啟 |
AOF 重播所有資料 — cache、lock、計數器 — 保留原本的到期時間。過期 lock 在 7s 內自動過期(與 leader 崩潰相同);cache 維持在 5 分鐘的有效窗口內;計數器繼續計數。最多 ~1s 的寫入可能遺失(appendfsync everysec 預設值)。 |
重啟期間短暫不可用(幾秒鐘);AOF 重播後恢復正常讀取 | 0 次 — cache 和計數器都存活;不需要額外取資料 | 透過 AOF + persistent volume 自我修復。1,000/天上限跨重啟仍然成立,不只是在單次運行期間。 |
| Leader 在取資料中途崩潰 例如 kill -9、OOM、容器重啟 |
ensure 不會執行 → lock 保持持有直到 7s 自動過期 → followers 等待並在 6s 時超時。每個 leader 持有唯一的 UUID(owner id,非 fencing token),因此一個在過期後才醒來的緩慢 leader 不會意外刪除新 leader 的 lock。 |
504 "Request timed out waiting for upstream result" | ≤1 次 — leader 在打 upstream 前崩潰為 0 次,期間或之後崩潰為 1 次 | 等待最多 7s 讓 lock 自動過期(即 lock TTL)。之後下一個請求成為新的 leader。 |
| Upstream 逾時 5s HTTParty timeout(~7%) |
RateApiClient 捕捉 Net::ReadTimeout,拋出 TimeoutError → 釋放 lock |
504 "Upstream timed out" | 1 次 — 失敗的嘗試消耗了配額 | 下一個請求成為新的 leader 並重試 upstream。 |
| Upstream HTTP 5xx ~7% 機率,來自真實測試 |
RateApiClient 看到非 200 狀態,拋出 UpstreamError → 釋放 lock,不寫入 cache | 502 "Upstream error" | 1 次 — 失敗的嘗試消耗了配額 | 下一個請求成為新的 leader 並重試 upstream。 |
| Upstream 回傳 200 但 body 為錯誤 ~7% 機率,來自真實測試 |
RateApiClient 解析 body,發現沒有 rate 欄位,拋出 UpstreamError → 釋放 lock,不寫入 cache |
502 "Upstream error" | 消耗 1 次 — 但我們刻意不 cache 錯誤回應。下一個請求重試,cache 中不會留下過期的垃圾資料。 | 下一個請求成為新的 leader 並重試 upstream。 |
| Upstream 回傳垃圾值 未知飯店、零或負數房價、小數字串、缺少欄位 |
RateApiClient 對每筆記錄執行 dry-schema + 正數檢查。壞記錄被記錄並跳過;其餘仍然被 cache。 | 若請求的房價通過驗證,服務仍回傳該值;否則回傳 502 | 1 次 — 部分回應仍勝於完全沒有 | 一筆壞資料不會讓 35 筆好資料也無法被 cache。如果所有記錄都無效,整個回應才會被拒絕。 |
| Upstream 401(驗證失敗) 例如需要更換 token |
RateApiClient 拋出 UnauthorizedError → PricingService 捕捉 → 退還配額單位(呼叫在處理前就被拒絕) | 502 "Upstream authentication failed" | 1 次但配額已退還 — QuotaGuard.refund_quota! 確保無效 token 不會消耗預算 |
操作人員必須更換 token。在此之前,每次 cache miss 消耗 1 個配額然後退還 — 預算保持完整。 |
| Upstream HTTP 429(他們的 rate limit) ~0.1% 機率 — upstream 每日上限已達 |
RateApiClient 拋出 QuotaExhaustedError → PricingService 捕捉 | 429 + Retry-After header |
1 次 — 被拒絕的呼叫在 upstream 端仍計入 | 等待 upstream 的窗口重置(假設 UTC 午夜,見決策 4)。我們的 QuotaGuard 在呼叫前就在本地捕捉相同條件 — 這一行是我們的計數器與 upstream 計數器偶爾出現差異的罕見競態。 |
| 配額計數器達到 1,000 當日預算耗盡 |
QuotaGuard 的原子 INCR 回傳 1001 → 在 RateApiClient 被呼叫之前拋出 ExhaustedError | 429 + Retry-After header(到 UTC 午夜的秒數) |
0 次 — 確定性保證 1,000 永遠不會被超出(已由壓力與混沌測試中的配額硬性上限測試驗證) | 計數器透過 EXPIREAT 在 UTC 午夜自動重置。命中 cache 的使用者仍能取得房價 — 只有需要重新取資料的請求會被阻擋。 |
每個依賴失敗都導向 5xx 回應。當保護層(cache / lock / quota)降級時,服務絕不打 upstream。這就是「fail-closed」— 以自身可用性換取保護 upstream 合約,對 proxy 來說這是正確的優先順序。
實作前對 upstream rate API 進行兩次真實測試 — 各 1,001 次呼叫。~28% 失敗率,涵蓋 4 種不同的失敗模式,在單次與批次請求中均已確認。
rates key(~7%)、回應中悄悄缺少 rate 欄位(~7%)、rates 陣列為空。光靠 HTTP 狀態碼無法信任。RateApiClient 對每個 200 都檢查 response body — 每種失敗模式都拋出 UpstreamError,永不悄悄回傳 nil。rate 欄位不等於配額耗盡rate 欄位代表配額耗盡。測試證明並非如此 — 缺少欄位是間歇性處理錯誤(~7%,與配額無關)。配額耗盡有自己的訊號:429 加上 {"error":"Rate limit exceeded (1000/day)"}。QuotaExhaustedError。缺少 rate → UpstreamError。兩種獨立例外,兩條獨立回應路徑。QuotaGuard 在每次 upstream 呼叫前預先消耗一個單位。發生 401 時退還該單位 — 請求從未到達 upstream 的處理階段。Unit test 驗證程式碼正確性,壓力測試驗證真實流量下的行為。我對真實運行的應用程式(Rails + Redis + rate-api)執行了 4 項測試 — 以下是每項測試的問題與量測結果。
| 測試項目 | 驗證問題 | 量測結果 | |
|---|---|---|---|
| Single-flight lock | 50 個使用者同時在 cold cache 狀態下打 API,是全部各自打 upstream,還是 lock 將它們合併成 1 次呼叫? | 1,404 個客戶端請求 → 1 次 upstream 呼叫 50 並發連線持續 5 秒 |
✓ PASS |
| Cache 命中率 | 在穩定流量下,有多少比例的請求從 cache 回傳,有多少打到 upstream? | 99.93% 命中率 30 秒內 7,081 次命中 / 5 次未命中 |
✓ PASS |
| 配額硬性上限 | 若出錯而接近每日 1,000 次上限,第 1,001 次呼叫會被阻擋,還是會漏過去? | 回傳 HTTP 429,0 次 upstream 呼叫 把計數器設為 1000,下一個請求被擋下並帶 Retry-After: 31254s(到 UTC 午夜) |
✓ PASS |
| Redis 崩潰時 fail-closed | Redis 崩潰時,服務是繼續意外打 upstream,還是乾淨地關閉入口? | 5/5 請求被擋下,0 次 upstream 呼叫 全部回傳 HTTP 504;Redis 恢復後服務自動回復 |
✓ PASS |
自行驗證:cd load_test && ./run_all.sh(約 6 分鐘)或 DURATION=30s ./run_all.sh(約 1 分鐘快速驗證)。
A Case Study in Caching & Quota Design
An application-layer cache gateway that minimizes real calls to an inference API under a hard request budget.
Every upstream call burns expensive inference compute. So I scoped the goal beyond “stay under 1,000/day”: minimise real calls. Our service must serve 1,000,000+ requests/day from this 1,000/day budget.
36 parameter combinations (4 periods × 3 hotels × 3 rooms) × 288 five-minute windows/day:
Way over the 1,000-call budget. We need batch fetching + request coalescing + a hard quota cap.
The full picture: client → Rails app (4 layers) → Redis + upstream. Each layer is detailed in its own section below.
This implementation uses a single Redis to hold the lock. An honest limit up front: in rare cases two leaders can exist at once (both requests call upstream). There are two separate causes:
Cause 1: the leader stalls (happens even on a single Redis)
The lock auto-expires after 7s (so a dead leader can't hold it forever). But if the leader stalls for more than 7s mid-work (e.g. a Ruby GC pause, or slow upstream), the lock expires first → the next request thinks nobody is working → grabs the lock → becomes a second leader.
Cause 2: Redis fails over (only with multiple nodes)
If you later move to Sentinel / Cluster, the moment a primary dies and a replica takes over, the lock write may not have replicated yet → the replica has no lock → a second leader appears.
What fixes each?
Cause 2 → a consensus-backed coordinator (etcd / Consul) guarantees the lock survives failover.
Cause 1 → consensus can't fix it: a stalled request wakes up still thinking it holds the lock. And note — even a fencing token only stops the stale leader's write from corrupting the resource; it does not undo the duplicate upstream call that already went out (preventing that needs idempotency, with upstream support).
Our choice: don't chase "exactly one leader."
A duplicate upstream call is harmless here (same data). The real backstop is the next layer — QuotaGuard's atomic counter hard-caps the daily budget at 1000, so even if two leaders slip through, the daily total can never be exceeded. The lock is an efficiency measure, not a correctness one; correctness lives in QuotaGuard.
Aside: Redlock (grab a majority across several Redis nodes) improves single-node fault tolerance, but still has known issues under clock skew or GC pauses — see Martin Kleppmann's distributed-locking analysis.
app/services/rate_cache.rb
The upstream API is expensive (1,000 calls/day). The same (period, hotel, room) returns the same rate for a while.
Without caching, every request triggers an upstream call. At 1,000,000 requests/day we'd blow through the quota very fast.
Cache each rate by (period, hotel, room) with a 5-minute TTL. Within that window, all requests for the same combo are served from cache — zero upstream calls.
cache key: "pricing:Summer:FloatingPointResort:SingletonRoom" TTL: 5 minutes Total keys: 36 (4 periods × 3 hotels × 3 rooms)
app/services/lock_and_fetch_rate.rb
Cache entries expire every 5 minutes. At the moment of expiry, many requests can arrive at the same time and all see a cache miss.
Without coordination, every concurrent cache-miss request calls upstream independently. A burst of 50 requests burns 50 quota units instead of 1.
Use a Redis lock (SET NX) with a global lock key (lock_and_fetch_rate:pricing:bulk_fetch) to elect one “leader”. The leader calls upstream and fills all 36 cache entries. All “followers” wait and read from the freshly populated cache.
Three values form a chain — each depends on the previous:
Without a timeout, if upstream hangs, the request blocks forever and ties up a Puma worker. 5s gives upstream enough time to respond normally, but frees the worker quickly when upstream is down.
The leader could get the upstream response at 4.9s, write cache, but the follower already gave up at 5.0s — wasted the whole wait. The 1s buffer covers this edge case.
Normally ensure deletes the lock right away after the leader finishes. But if the process crashes hard (e.g. kill -9), ensure never runs and the lock is stuck. 7s auto-expire is the safety net.
It must be longer than FOLLOWER_TIMEOUT (6s). If shorter, the lock expires while followers are still waiting → a new request grabs the lock → becomes a second leader → calls upstream again → wastes quota.
1ms = too fast, floods Redis. 1s = too slow, follower waits up to 1s after leader finishes. 50ms = good balance. At most 50ms extra latency.
app/services/quota_guard.rb
With Batch Fetch, the theoretical maximum is only 288 calls/day — well under the 1,000 budget. Under normal operation, QuotaGuard will never trigger. It exists as a local tripwire for scenarios that shouldn't happen but could — bugs that turn every cache miss into an upstream call:
rate to price. Parser fails, cache stays empty, every request hits upstream until upstream itself starts rejecting.RateCache.write_all has a typo in the key name. Writes succeed silently but reads never match → same runaway loop.Normally the 1,001st call is rejected up front with 429 — upstream never runs the expensive inference for it. But if that limiter ever misconfigures (bug, redeploy, staging quirk), our overflow would actually hit the computationally expensive inference and burn real compute. QuotaGuard keeps us honest to the 1,000/day contract regardless of upstream's enforcement state.
Trade-offs accepted:
DAILY_LIMIT needs updating too.Multiple Puma workers handle requests in parallel. A naive “read → check → increment” has a race condition: two workers could both read 999, both pass the check, and both call upstream.
// Race condition example: Worker A: read count → 999 → passes check → calls upstream // count is now 1000 Worker B: read count → 999 → passes check → calls upstream // EXCEEDED!
Redis INCR is atomic — it increments and returns the new value in one step. Each request gets a unique sequential number. Exactly 1,000 calls pass through. No race condition.
daily key: "quota:2026-04-22" // auto-expires at midnight Redis INCR → 1000 → pass ✓ Redis INCR → 1001 → raise ExhaustedError ✗
Because INCR is atomic, there's no way to overshoot. A lower buffer only makes sense with a gradual slowdown strategy. We use a hard stop with Retry-After until midnight, so capping lower just wastes quota.
The upstream the upstream rate API documentation does not specify when the daily quota resets. This implementation assumes UTC midnight — the Redis key uses quota:YYYY-MM-DD (UTC) and the Retry-After header counts seconds until UTC 00:00. If the actual reset time differs, only the key date logic and Retry-After calculation need to change.
lib/rate_api_client.rb → get_all_rates
The upstream API accepts multiple rate queries in a single request. There are 36 possible combinations (4 periods × 3 hotels × 3 rooms).
Without batching, each cache miss fetches only one rate — one quota unit per rate. Even with LockAndFetchRate, the worst case is 36 cache misses × 288 windows = 10,368 calls/day.
On any cache miss, fetch all 36 rates in a single API call via RateApiClient.get_all_rates. Then RateCache.write_all fills every cache key at once. One quota unit refreshes the entire cache for the next 5 minutes.
Cache miss for "Summer:FloatingPointResort:SingletonRoom" → fetch 1 rate → write 1 cache entry → 1 quota unit for 1 rate
Worst case: 10,368 calls/day
Cache miss for "Summer:FloatingPointResort:SingletonRoom" → fetch 36 rates (all combos) → write 36 cache entries → 1 quota unit for 36 rates
Worst case: 288 calls/day (= theoretical minimum)
Each decision lists the options considered, the choice, and the trade-off consciously accepted. Click to expand.
Bulk fetching itself isn't a real choice — the math forces it: 36 combos × 288 cache windows/day = 10,368 calls, way over the 1,000 budget. So each upstream call has to return all 36 rates. Under a 5-min TTL, 288/day is the math floor on the upper bound — no compliant design can do fewer at peak. The interesting question is when to trigger that call: on the first cache miss (lazy — 0 at the lower bound), or on a schedule (prewarming — always ≥ 288).
| Option | Behaviour |
|---|---|
| Periodic prewarming e.g. Sidekiq cron every 4 min |
Cache always warm, every request is instant. But: burns 360 calls/day even at zero traffic; needs extra infra (Sidekiq + scheduler). |
| Lazy on-demand ✓ | Upstream cost proportional to actual demand — 0 calls when nobody's asking, up to 288/day at full traffic. No extra infra. |
Why lazy fits this workload:
Prewarming is the right call when: traffic is predictable and steady, latency SLAs are strict (any cache miss is unacceptable), or the quota budget is generous. For those workloads, prewarming wins. Different constraint, different answer.
Trade-off accepted: the first user to hit an expired cache pays a small latency tax (waits for the upstream call). All subsequent requests in that 5-min window are instant cache hits.
| Option | Rejected because |
|---|---|
MemoryStore | Per-Puma-worker memory — hit ratio divided by worker count, breaks 10:1 requirement. |
SolidCache | SQLite-based; no atomic INCR for quota; row-level lock contention. |
| Redis ✓ | Cross-worker sharing; atomic primitives (SET NX, INCR, EXPIREAT) solve cache, lock, and quota with one dependency. |
Trade-off: one extra service in docker-compose.yml. Worth it — Redis is ubiquitous in Rails production.
Without coalescing, a 50-request burst burns 50 units of the daily budget in milliseconds. Redis SET NX elects one leader across all Puma workers; followers poll the cache.
// LockAndFetchRate pseudocode: token = SecureRandom.uuid // per-leader owner id (safe release, NOT a fencing token) if SET NX "lock:bulk_fetch" = token (TTL 7s) → leader yield // fetch upstream, write cache Lua: if GET == token then DEL // release only if still ours else → follower: poll cache every 50ms up to 6s
Owner id (safe release): each leader holds a unique UUID. On release, a Lua script checks "still mine?" before deleting — otherwise a slow leader could wipe the new leader's lock and let a third request slip in. Note: this is an owner id, not a fencing token — it only prevents deleting the wrong lock; it does not stop a stale leader from making a duplicate upstream call after the lock expires (that needs a fencing token + resource check, or idempotency).
Single-Redis assumption: two leaders can briefly coexist, from two causes: (1) a leader stalling (GC pause, or upstream slower than the lock TTL) so the lock expires — this happens even on a single Redis; (2) a lost lock write during primary failover under Sentinel / Cluster. Raft-based coordinators (etcd / Consul) fix only cause 2, not cause 1 — the only fix for that is a fencing token validated by the protected resource. We deliberately don't: the lock is an efficiency measure, and correctness comes from QuotaGuard's atomic counter (fail-closed) — even if two leaders co-exist briefly, the daily cap still holds.
The upstream limit is 1,000 calls per day — not per second. So we need a simple counter that resets at midnight.
Batch Fetch already keeps the theoretical max at 288 calls/day — well under budget. Under normal operation, QuotaGuard never triggers.
The honest remaining value is defensive client behavior: we cap our own call rate at 1,000/day even if upstream's own rate limiter fails to enforce it (bug, redeploy, staging quirk). In that edge case our overflow would otherwise hit upstream's computationally expensive inference and burn real compute. QuotaGuard prevents that.
Trade-offs accepted:
DAILY_LIMIT needs updating too.The counter must be shared across all Puma workers (separate processes), and INCR is atomic — it increments and returns the new value in one step, so two workers can never both read 999 and both pass the check.
INCR on quota:YYYY-MM-DD — atomic, cross-workerINCR is atomic — no way to overshootrefund_quota! wraps DECR in a Lua script that only runs when counter is above zero. Without this, burst 401s before any INCR would push counter negative, letting subsequent INCRs silently exceed 1,000.Quota resets at UTC midnight. The upstream docs don't specify when the daily limit resets, so we assume UTC 00:00. The Redis key uses quota:YYYY-MM-DD (UTC) and Retry-After counts seconds until UTC midnight. If the actual reset time differs, only these two values need updating.
When a dependency dies (Redis, upstream, lock), what should the service do? There are four common strategies — each with a different priority.
| Strategy | What it does | Protects | Sacrifices |
|---|---|---|---|
| ✓ Fail-closed My choice |
Dependency dies → return 5xx, never call upstream | Upstream contract (1000/day cap) | Own service uptime |
| Fail-open | Dependency dies → bypass protection, hit upstream directly | Own uptime | Quota burns fast; thundering herd unblocked |
| Graceful degradation | Return stale cache, default value, or partial response | User experience | Data freshness / correctness |
| Fail-fast + retry | Return 5xx, expect client to retry with backoff | Same as fail-closed but pushes retry to caller | Assumes well-behaved client |
Why fail-closed for this service: two spec rules forced it.
Once those two are eliminated, fail-closed is the only honest choice. Fail-fast + retry is functionally equivalent but offloads correctness to the client; for a proxy, owning that responsibility is cleaner.
When I'd choose differently:
The right choice always depends on whether "wrong data" or "no data" is more harmful. For pricing, wrong data wins.
Every upstream failure — 401, 5xx, timeout — returned 400 Bad Request. The client had no way to distinguish a bad request from a server-side failure.
Each failure maps to a distinct HTTP status. The client knows exactly what happened and whether to fix the request, retry, or escalate.
{"message":"Failed to process rates...","status":"error"}rate field missing from responseFrom 1,001 live test calls — see UPSTREAM_BEHAVIOR.md for full pattern tables.
Anti-Corruption Layer (ACL) RateApiClient absorbs all upstream quirks and exposes a clean typed interface. PricingService never touches a raw HTTP response:
Integer on success| Situation | Who raises | Client status | Client message | Client action |
|---|---|---|---|---|
| Invalid params | Controller | 400 | Missing required parameters / Invalid period | Fix request |
| Daily quota used up | QuotaGuard | 429 | Daily quota reached — requests resume at midnight UTC | Wait for Retry-After |
| Upstream returns HTTP 429 (token limit) | RateApiClient | 429 | Upstream API quota exhausted — token limit reached | Wait for Retry-After |
| Bad token (401) | RateApiClient | 502 | Upstream authentication failed | — |
| Upstream 200 with error body | RateApiClient | 502 | Upstream error | — |
| Upstream 200, rate field absent | RateApiClient | 502 | Upstream error | — |
| Upstream 5xx | RateApiClient | 502 | Upstream error | — |
| Upstream timeout | RateApiClient | 504 | Upstream timed out | — |
| Follower waited too long | LockAndFetchRate | 504 | Request timed out waiting for upstream result | — |
| Unexpected error | StandardError catch-all | 500 | Unexpected error | — |
Both 429 responses include a Retry-After header (seconds until UTC midnight). All other error responses return only the status code and message — no retry guidance is provided.
Every request emits a single JSON log line with: method, path, status, duration, cache_status ("hit" or "miss"), and upstream_latency_ms.
// Example log line: {"method":"GET","path":"/api/v1/pricing","status":200,"duration":12.3, "cache_status":"hit","upstream_latency_ms":null}
Why Lograge over paid services (Datadog, New Relic): Lograge is free, zero config, and ships with Rails. For a case study like this, structured JSON logs are enough to prove the observability mindset. In production, these same JSON lines can be piped into any log aggregator.
Why this matters: "How would you know this service is healthy in production?" has a concrete answer — filter by cache_status=miss, watch upstream_latency_ms.
All test files follow a consistent section structure: Happy path → Fail path → Edge/Boundary → Isolation. Coverage measured with SimpleCov — currently 100% (scaffold files excluded).
Why WebMock over .stub: Stubbing RateApiClient.get_all_rates directly bypasses the HTTP layer, hiding timeout, header, and retry behaviour. WebMock intercepts at the socket layer, producing more realistic tests.
What happens when each dependency fails, and why the 1000/day cap still holds.
| What fails | What happens inside | Client sees | Upstream impact | Recovery |
|---|---|---|---|---|
| Redis fully down e.g. crash, OOM, network partition |
Rails' built-in cache error handler swallows the Redis error and returns nil → Rails.cache.write(..., unless_exist: true) returns nil → LockAndFetchRate falls into the follower branch → poll_for_result sees nil for 6s → raises LockAndFetchRate::TimeoutError |
504 "Request timed out waiting for upstream result" | 0 calls — no one acquires the lock, so upstream is never reached | Service recovers immediately when Redis comes back online (verified by the Fail-closed when Redis dies test in Load & Chaos Tests) |
| Redis restart e.g. deploy, OOM, container restart |
AOF replays everything back — cache, lock, counter — with their original expiry times. Stale lock self-expires within 7s (same as leader crash); cache stays inside the 5-min freshness window; counter resumes counting. Up to ~1s of writes may be lost (appendfsync everysec default). |
Brief unavailability during restart (a few seconds); reads resume normally after AOF replay | 0 calls — cache and counter both survive; nothing extra to fetch | Self-healing via AOF + persistent volume. The 1000/day cap holds across restarts, not just within a single uptime. |
| Leader crashes mid-fetch e.g. kill -9, OOM, container restart |
ensure never runs → lock stays held until 7s auto-expire → followers wait then time out at 6s. Each leader holds a unique UUID (owner id, not a fencing token), so a slow leader waking up after expiry won't accidentally delete the new leader's lock. |
504 "Request timed out waiting for upstream result" | ≤1 call — 0 if leader crashed before calling upstream, 1 if it crashed during or after | Wait up to 7s for the lock to auto-expire (that's the lock TTL). Then the next request takes over as new leader. |
| Upstream timeout 5s HTTPParty timeout (~7%) |
RateApiClient catches Net::ReadTimeout, raises TimeoutError → lock released |
504 "Upstream timed out" | 1 call — quota consumed for the failed attempt | Next request becomes new leader and retries upstream. |
| Upstream HTTP 5xx ~7% rate observed in live testing |
RateApiClient sees non-200 status, raises UpstreamError → lock released, cache not written | 502 "Upstream error" | 1 call — quota consumed for the failed attempt | Next request becomes new leader and retries upstream. |
| Upstream returns 200 with error body ~7% rate observed in live testing |
RateApiClient parses body, sees no rate field, raises UpstreamError → lock released, cache not written |
502 "Upstream error" | 1 call burned — but we deliberately don't cache the bad response. Next request retries, no stale garbage stuck in cache. | Next request becomes new leader and retries upstream. |
| Upstream returns garbage values unknown hotel, zero/negative rate, decimal strings, missing fields |
RateApiClient runs each record through a dry-schema + positive-numeric check. Bad records are logged and skipped; the rest still get cached. | Service still returns the requested rate if it survived validation, otherwise 502 | 1 call — partial response still better than zero | One bad row doesn't deny cache to 35 good rows. If every record is invalid the whole response is rejected. |
| Upstream 401 (auth failure) e.g. token rotation needed |
RateApiClient raises UnauthorizedError → PricingService catches → refunds the quota unit (the call was rejected before processing) | 502 "Upstream authentication failed" | 1 call but quota refunded — QuotaGuard.refund_quota! ensures bad tokens don't burn budget |
Operator must rotate token. Until then, every cache miss costs 1 quota then refunds — budget stays intact. |
| Upstream HTTP 429 (their rate limit) ~0.1% observed — upstream's daily limit hit |
RateApiClient raises QuotaExhaustedError → PricingService catches | 429 + Retry-After header |
1 call — rejected call still counts upstream-side | Wait until upstream's window resets (assumed UTC midnight, see D4). Our QuotaGuard catches the same condition locally before the call — this row is the rare race where our counter and upstream's diverge. |
| Quota counter at 1000 budget exhausted for the day |
QuotaGuard's atomic INCR returns 1001 → raises ExhaustedError before RateApiClient is called | 429 + Retry-After header (seconds until UTC midnight) |
0 calls — the deterministic guarantee that 1000 is never exceeded (verified by the Quota hard cap test in Load & Chaos Tests) | Counter auto-resets at UTC midnight via EXPIREAT. Users hitting cache still get their rate — only fresh fetches get blocked. |
Every dependency failure routes to a 5xx response. The service never calls upstream when its protection layers (cache / lock / quota) are degraded. This is "fail-closed" — protect the upstream contract at the cost of own uptime, which is the right priority for a proxy.
Two live test runs — 1001 calls each — against the upstream rate API, run before implementation. ~28% failure rate across 4 distinct failure modes, all confirmed in both single and bulk requests.
rates key (~7%), rate field silently absent from response (~7%), and empty rates array. HTTP status alone cannot be trusted.RateApiClient inspects the response body on every 200 — each failure mode raises UpstreamError, never returns nil silently.rate field is not quota exhaustionrate field meant quota was exhausted. Testing proved otherwise — the missing field is an intermittent processing error (~7% regardless of quota). Quota exhaustion has its own signal: 429 with {"error":"Rate limit exceeded (1000/day)"}.QuotaExhaustedError. Map missing rate → UpstreamError. Two separate exceptions, two separate response paths.QuotaGuard pre-consumes one unit before every upstream call. On 401, it refunds the unit — the request never reached upstream processing.Unit tests prove the code is correct. Load tests prove it works under real traffic. I ran 4 tests against the live app (Rails + Redis + rate-api) — here's what each one asks and what I measured.
| What I'm testing | The question | What I measured | |
|---|---|---|---|
| Single-flight lock | If 50 users hit the API at the same time on a cold cache, do they all hammer the upstream — or does my lock collapse them into one call? | 1,404 client requests → 1 upstream call 5 seconds at 50 concurrent connections |
✓ PASS |
| Cache hit rate | Under steady traffic, what percent of requests are served from cache vs. hitting the upstream? | 99.93% hit rate 7,081 hits / 5 misses in 30 seconds |
✓ PASS |
| Quota hard cap | If something goes wrong and we approach the 1,000/day upstream limit, does the 1,001st call get blocked — or does it leak through? | HTTP 429 returned, 0 upstream calls Set counter to 1000, next request blocked with Retry-After: 31254s (until UTC midnight) |
✓ PASS |
| Fail-closed when Redis dies | If Redis crashes, does the service keep accidentally calling upstream — or does it shut the door cleanly? | 5/5 requests blocked, 0 upstream calls All returned HTTP 504; service auto-recovered when Redis came back |
✓ PASS |
Try it yourself: cd load_test && ./run_all.sh (~6 minutes total) or DURATION=30s ./run_all.sh (~1 minute smoke test).