Repository navigation
[API] Fix: intermittent login 401 on high-latency links (Node happy-eyeballs 250ms budget) - #381
Merged
Merged
Conversation
Node >=20 enables autoSelectFamily with a 250ms per-address connect budget. On links whose TCP connect to the identity provider exceeds it (VPNs, deployments far from Microsoft endpoints), every outbound connection attempt is aborted: jwks-rsa cannot download signing keys and the API rejects ALL valid tokens with an empty 401 — intermittently, since real-world latency fluctuates around the threshold. The same budget governs the http/https agents used by the storage SDKs. Known upstream issue (nodejs/node#54359). Node 25.2 raised the default to 500ms (nodejs/node#60334), still tight for high-RTT links and not backported to 24 LTS, so set the 2500ms value originally proposed upstream (nodejs/node#56738) at bootstrap. Connections that finish faster are unaffected; only failover to dead addresses gets slower.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What & why
autoSelectFamily) with a 250ms per-address connect budget. Above it, every address attempt aborts and the connection fails with an empty-messageAggregateError. Known upstream issue: nodejs/node#54359 (also nodejs/node#52216, nodejs/undici#2777).ciamlogin.com), which makes it look like an auth bug and very hard to diagnose — the API only logs{"code":"UNAUTHORIZED","message":""}.Approach
Raise the per-address budget process-wide at API bootstrap via
net.setDefaultAutoSelectFamilyAttemptTimeout. The budget lives innet.connect, so the same 250ms limit governs both thehttp/httpsagents (jwks-rsa, Azure Blob SDK, AWS S3 SDK) and nativefetch(undici reads net's defaults at connect time). A fix scoped to the JWKS client (jwks-rsa supportsrequestAgent/fetcher) was rejected: it would leave the storage SDKs broken on the same links.Key changes & decisions
1. Widen the Happy Eyeballs connect budget at bootstrap
apps/api/src/app.ts,apps/api/src/config/constants.ts· 5c0e343JwksClient.getSigningKeydying on Node's 250ms-per-address connect budget.net.setDefaultAutoSelectFamilyAttemptTimeout(NETWORK_CONNECTION_ATTEMPT_TIMEOUT_MS)(2500ms) at the top ofapp.ts; constant documented inconfig/constants.tsper the repo's deployment-tunable convention.--no-network-family-autoselection) was rejected: it removes the IPv6→IPv4 failover that the feature gets right.NODE_OPTIONSin compose/Dockerfile was rejected because it silently misses local dev — where this was hit. Raising the budget keeps the failover; the only cost is slower failover to genuinely dead addresses.Prior art & references — this exact mitigation is established practice
(every reference below was verified by fetching the source)
The author of Node's Happy Eyeballs implementation prescribes this exact fix. In nodejs/node#54359, @ShogunPanda (Paolo Insogna, who implemented
autoSelectFamily) responds to the production reports:A bootstrap-time call to this setter is, verbatim, the supported remedy.
Node core made the same change to its own default. Node v25.2.0 raised the default from 250ms to 500ms (nodejs/node#60334, authored by Rod Vagg/TSC, ~10 maintainer approvals): "The current 250ms timeout is too short for high-latency network environments." Not backported to Node 24 LTS — our runtime keeps 250ms, hence this PR.
AWS SDK for JavaScript v3 — official guidance. AWS's Effective Practices doc, section "(4) Allow more time to establish connections when making requests cross-region" describes the same symptom (
AggregateError/ETIMEDOUT), cites the same issue, and recommends the same flag ornet.setDefaultAutoSelectFamilyAttemptTimeout(500)in application code: "The default value of 250ms may be too low … or simply in conditions of low network speed."It is a documented, sanctioned knob. Node ships a dedicated CLI flag (
--network-family-autoselection-attempt-timeout, since 20.13/22.1) and documentsnet.setDefaultAutoSelectFamilyAttemptTimeout.Well-known products ship the same mitigation in their own code:
packages/cli/bin/n8n, shipped since n8n PR #9860):require('net').setDefaultAutoSelectFamily?.(false)— structurally identical bootstrap pattern, opting to disable rather than raise.autoSelectFamily/autoSelectFamilyAttemptTimeoutas first-classMongoClientOptions(since v6.9.0).NODE_OPTIONS="--network-family-autoselection-attempt-timeout=500"for this exact error, citing the same Node issue.autoSelectFamilyAttemptTimeout = 5000insrc/main/network/http.tsafter users on high-latency islands hit ETIMEDOUT (thorium-reader#3444); LiveKit agents-js docs recommend--network-family-autoselection-attempt-timeout=5000for the same symptom.On the value: our 2500ms sits between Node's new 500ms default (tuned for general use) and the 5000ms shipped by Thorium Reader / recommended by LiveKit — chosen because measured VPN connect already spiked to 430ms and target deployments are countries on potentially slow links. It only bounds connection-establishment attempts; request/socket/keep-alive timeouts are untouched.
Test plan
pnpm format/pnpm lint/pnpm type-check— green.pnpm test --filter=api -- /getOrganizationById/integration.test.ts— app boots and serves with the setter in place (26/26).fetchandJwksClient.getSigningKeyagainst the tenant JWKS failed withAggregateError; re-run under--network-family-autoselection-attempt-timeout=2000(the flag equivalent of this setter) → 200 and key resolution OK.Risk / blast radius