Skip to content

[API] Fix: intermittent login 401 on high-latency links (Node happy-eyeballs 250ms budget) - #381

Merged
mrivas00 merged 1 commit into
mainfrom
fix/mati/net-autoselect-attempt-timeout
Jun 12, 2026
Merged

mrivas00 merged 1 commit into
mainfrom
fix/mati/net-autoselect-attempt-timeout

Conversation

@mrivas00

@mrivas00 mrivas00 commented Jun 12, 2026 •

Copy link
Copy Markdown
Collaborator

What & why

  • On links where TCP connect to the identity provider exceeds 250ms (VPNs, deployments far from Microsoft endpoints), every login fails with an empty 401 — the API cannot download JWKS signing keys, so all valid tokens are rejected.
  • Root cause is not ours: Node ≥20 enables Happy Eyeballs (autoSelectFamily) with a 250ms per-address connect budget. Above it, every address attempt aborts and the connection fails with an empty-message AggregateError. Known upstream issue: nodejs/node#54359 (also nodejs/node#52216, nodejs/undici#2777).
  • The failure is intermittent because real-world latency fluctuates around the threshold (measured 190–430ms over a VPN against ciamlogin.com), which makes it look like an auth bug and very hard to diagnose — the API only logs {"code":"UNAUTHORIZED","message":""}.

Approach

Raise the per-address budget process-wide at API bootstrap via net.setDefaultAutoSelectFamilyAttemptTimeout. The budget lives in net.connect, so the same 250ms limit governs both the http/https agents (jwks-rsa, Azure Blob SDK, AWS S3 SDK) and native fetch (undici reads net's defaults at connect time). A fix scoped to the JWKS client (jwks-rsa supports requestAgent/fetcher) was rejected: it would leave the storage SDKs broken on the same links.

Key changes & decisions

1. Widen the Happy Eyeballs connect budget at bootstrap

apps/api/src/app.ts, apps/api/src/config/constants.ts · 5c0e343

  • Trigger — logins 401 behind a VPN with a perfectly valid token; reproduced down to JwksClient.getSigningKey dying on Node's 250ms-per-address connect budget.
  • Change — net.setDefaultAutoSelectFamilyAttemptTimeout(NETWORK_CONNECTION_ATTEMPT_TIMEOUT_MS) (2500ms) at the top of app.ts; constant documented in config/constants.ts per the repo's deployment-tunable convention.
  • Why this way — this is the remedy the Node maintainers themselves prescribe (see Prior art below). Disabling autoselection entirely (--no-network-family-autoselection) was rejected: it removes the IPv6→IPv4 failover that the feature gets right. NODE_OPTIONS in compose/Dockerfile was rejected because it silently misses local dev — where this was hit. Raising the budget keeps the failover; the only cost is slower failover to genuinely dead addresses.

Prior art & references — this exact mitigation is established practice

(every reference below was verified by fetching the source)

  1. The author of Node's Happy Eyeballs implementation prescribes this exact fix. In nodejs/node#54359, @ShogunPanda (Paolo Insogna, who implemented autoSelectFamily) responds to the production reports:

    "You can use --network-family-autoselection-attempt-timeout or net.setDefaultAutoSelectFamilyAttemptTimeout to adjust the default value." … "since all you need is to use a simple API call when booting your process, I think we are fine the way we are now."

    A bootstrap-time call to this setter is, verbatim, the supported remedy.

  2. Node core made the same change to its own default. Node v25.2.0 raised the default from 250ms to 500ms (nodejs/node#60334, authored by Rod Vagg/TSC, ~10 maintainer approvals): "The current 250ms timeout is too short for high-latency network environments." Not backported to Node 24 LTS — our runtime keeps 250ms, hence this PR.

  3. AWS SDK for JavaScript v3 — official guidance. AWS's Effective Practices doc, section "(4) Allow more time to establish connections when making requests cross-region" describes the same symptom (AggregateError/ETIMEDOUT), cites the same issue, and recommends the same flag or net.setDefaultAutoSelectFamilyAttemptTimeout(500) in application code: "The default value of 250ms may be too low … or simply in conditions of low network speed."

  4. It is a documented, sanctioned knob. Node ships a dedicated CLI flag (--network-family-autoselection-attempt-timeout, since 20.13/22.1) and documents net.setDefaultAutoSelectFamilyAttemptTimeout.

  5. Well-known products ship the same mitigation in their own code:

    • n8n calls the global setter first thing in its bin script (packages/cli/bin/n8n, shipped since n8n PR #9860): require('net').setDefaultAutoSelectFamily?.(false) — structurally identical bootstrap pattern, opting to disable rather than raise.
    • MongoDB official Node driver exposes both autoSelectFamily/autoSelectFamilyAttemptTimeout as first-class MongoClientOptions (since v6.9.0).
    • Coinbase CDP SDK README troubleshooting prescribes NODE_OPTIONS="--network-family-autoselection-attempt-timeout=500" for this exact error, citing the same Node issue.
    • Thorium Reader (EDRLab) ships autoSelectFamilyAttemptTimeout = 5000 in src/main/network/http.ts after users on high-latency islands hit ETIMEDOUT (thorium-reader#3444); LiveKit agents-js docs recommend --network-family-autoselection-attempt-timeout=5000 for the same symptom.
  6. On the value: our 2500ms sits between Node's new 500ms default (tuned for general use) and the 5000ms shipped by Thorium Reader / recommended by LiveKit — chosen because measured VPN connect already spiked to 430ms and target deployments are countries on potentially slow links. It only bounds connection-establishment attempts; request/socket/keep-alive timeouts are untouched.

Test plan

  • pnpm format / pnpm lint / pnpm type-check — green.
  • pnpm test --filter=api -- /getOrganizationById/integration.test.ts — app boots and serves with the setter in place (26/26).
  • Mechanism verified during diagnosis on a failing link (VPN, ~285ms connect): plain fetch and JwksClient.getSigningKey against the tenant JWKS failed with AggregateError; re-run under --network-family-autoselection-attempt-timeout=2000 (the flag equivalent of this setter) → 200 and key resolution OK.
  • Reviewer repro (optional, needs a slow link/VPN): on current main behind a >250ms-connect VPN, login → 401 with valid OTP; on this branch → login succeeds.

Risk / blast radius

  • Process-global but narrow in semantics: only lengthens the per-address budget during outbound connection establishment. Connections that complete under 250ms are unaffected; requests, sockets, and keep-alive timeouts unchanged. Worst case: failover to a dead address takes up to 2500ms per address instead of 250ms.
  • No API contract, schema, or dependency changes.

Node >=20 enables autoSelectFamily with a 250ms per-address connect
budget. On links whose TCP connect to the identity provider exceeds it
(VPNs, deployments far from Microsoft endpoints), every outbound
connection attempt is aborted: jwks-rsa cannot download signing keys and
the API rejects ALL valid tokens with an empty 401 — intermittently,
since real-world latency fluctuates around the threshold. The same
budget governs the http/https agents used by the storage SDKs.

Known upstream issue (nodejs/node#54359). Node 25.2 raised the default
to 500ms (nodejs/node#60334), still tight for high-RTT links and not
backported to 24 LTS, so set the 2500ms value originally proposed
upstream (nodejs/node#56738) at bootstrap. Connections that finish
faster are unaffected; only failover to dead addresses gets slower.
@mrivas00
mrivas00 merged commit b9594e3 into main Jun 12, 2026
6 checks passed
@mrivas00
mrivas00 deleted the fix/mati/net-autoselect-attempt-timeout branch June 12, 2026 12:48
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant