Website: byokrelay.com | Hosted relay: relay.byokrelay.com
Your users bring their own AI keys. byok-relay lets them use those keys straight from the browser — CORS handled, keys never in your code, costs on their bill.
Browser apps can't call api.openai.com or api.anthropic.com directly — CORS blocks them. The usual fix (a backend proxy) puts your users' keys — and your users' AI costs — on your tab. byok-relay flips this: each user gets a secure token; they store their own key; they pay for their own inference. You build the product.
Option A — Use our relay (zero setup):
https://relay.byokrelay.com
Free. Open CORS (any origin). Health check →
Option B — Self-host in 3 commands:
git clone https://github.com/avikalpg/byok-relay.git && cd byok-relay
echo "ENCRYPTION_SECRET=$(openssl rand -hex 32)" > .env
docker compose up -d # relay running at http://localhost:3000Or without Docker: npm install && npm start (requires Node 18+). Full quickstart →
Trust model: The managed relay holds the
ENCRYPTION_SECRET. All request bodies (prompts, conversation history) transit through it in plaintext on the way to AI providers. It is suitable for prototypes, demos, and development — not production apps with paying users or sensitive data. For production: self-host. See SECURITY.md for full data residency details.
If you're using a coding agent (Cursor, Claude Code, Copilot, Codex, etc.), install the skill and let it handle the integration:
npx skills add avikalpg/byok-relayOr point your agent directly at the skill file:
https://byokrelay.com/skill
Prompt: "Read the byok-relay skill at https://byokrelay.com/skill and integrate byok-relay into this project using the hosted relay at https://relay.byokrelay.com"
Browser byok-relay AI Provider
│ │ │
├─ POST /users ────────────►│ │
│◄─ { token } ─────────────┤ │
│ │ │
├─ POST /keys/anthropic ───►│ │
│ { key: "sk-ant-..." } │ (stored encrypted) │
│◄─ { ok: true } ──────────┤ │
│ │ │
├─ POST /relay/anthropic ──►│ │
│ x-relay-token: <token> ├─ (real key injected) ►│
│ { model, messages... } │ │
│◄─ streamed response ──────┤◄─ streamed response ──┤
The token (not the key) lives in the browser. The API key stays server-side, encrypted at rest with AES-256-GCM. In the individual flow, users register once; every request uses their key and is billed to their provider account. In the B2B flow, requests use the organization's registered key and provider billing account.
| byok-relay | OpenRouter | LiteLLM | |
|---|---|---|---|
| Who holds the API keys | Your users | OpenRouter | Your org |
| Who pays for AI usage | Your users | You (the dev) | You (the org) |
| BYOK for end users | ✅ | ❌ | ❌ |
| Browser-safe (CORS handled) | ✅ | ✅ | ❌ (needs backend) |
| Self-hosted | ✅ | ❌ | ✅ |
| Open source | ✅ Apache 2.0 | ❌ | ✅ |
| Model routing / fallbacks | ❌ | ✅ | ✅ |
Use OpenRouter or LiteLLM when you're paying for your users' AI and want routing + analytics. Use byok-relay when you want users to bring their own keys.
The easiest way to integrate byok-relay into a Vite/ESM browser app:
npm install @byok-relay/clientimport { createClient } from '@byok-relay/client'
const relay = createClient({
relayUrl: import.meta.env.VITE_RELAY_URL ?? 'https://relay.byokrelay.com', // or your self-hosted relay URL
})
// Your user enters their API key once
await relay.storeKey('openai', userApiKey)
// Then stream — no backend required
const text = await relay.streamChat({
provider: 'openai',
model: 'gpt-4o-mini',
messages: [{ role: 'user', content: 'Hello!' }],
onChunk: (delta) => console.log(delta),
})Works in browsers (localStorage default), Node.js (in-memory default), and any custom storage adapter. See packages/client/README.md for full API reference.
Option A — zero install with npx:
ENCRYPTION_SECRET=$(openssl rand -hex 32) npx byok-relayOption B — clone and run:
# 1. Clone and install
git clone https://github.com/avikalpg/byok-relay.git && cd byok-relay && npm install
# 2. Configure
echo "ENCRYPTION_SECRET=$(openssl rand -hex 32)" > .env
echo "ALLOWED_ORIGINS=http://localhost:5173" >> .env # replace with your browser app's origin
# Production only: restrict who can register users, and keep this shell variable for step 4.
# APP_SECRET=$(openssl rand -hex 32)
# echo "APP_SECRET=$APP_SECRET" >> .env
# 3. Start
npm start &
i=0; until curl -fsS http://localhost:3000/health >/dev/null; do i=$((i + 1)); [ "$i" -ge 30 ] && { echo "Relay did not become ready"; exit 1; }; sleep 1; done
# 4. Register a user and get a token
# Development, when APP_SECRET is not set:
TOKEN=$(curl -s -X POST http://localhost:3000/users \
-H "Content-Type: application/json" \
-d '{"app_id":"test"}' | node -e "let s=''; process.stdin.on('data', d => s += d).on('end', () => console.log(JSON.parse(s).token))")
# Production, when APP_SECRET is set:
# TOKEN=$(curl -s -X POST http://localhost:3000/users \
# -H "Content-Type: application/json" \
# -H "Authorization: Bearer $APP_SECRET" \
# -d '{"app_id":"test"}' | node -e "let s=''; process.stdin.on('data', d => s += d).on('end', () => console.log(JSON.parse(s).token))")
# 5. Store your Anthropic key
curl -X POST http://localhost:3000/keys/anthropic \
-H "Content-Type: application/json" \
-H "x-relay-token: $TOKEN" \
-d '{"key":"sk-ant-YOUR-KEY-HERE"}'
# 6. Relay a request — unified endpoint
curl -X POST http://localhost:3000/relay \
-H "Content-Type: application/json" \
-H "x-relay-token: $TOKEN" \
-d '{"model":"anthropic/claude-3-5-haiku","max_tokens":256,"messages":[{"role":"user","content":"Hello!"}]}'
# Or with streaming
curl -X POST http://localhost:3000/relay \
-H "Content-Type: application/json" \
-H "x-relay-token: $TOKEN" \
-d '{"model":"gpt-4o","stream":true,"messages":[{"role":"user","content":"Hello!"}]}'| Provider | Name | Notes |
|---|---|---|
| Anthropic | anthropic |
Claude models, SSE streaming |
| OpenAI | openai |
GPT models, SSE streaming |
google |
Gemini API (key in query param) | |
| Groq | groq |
Fast inference, OpenAI-compatible |
| OpenRouter | openrouter |
200+ models via one API |
| Mistral | mistral |
Mistral models |
| Any OpenAI-compatible | openai-compatible |
Pass x-relay-base-url header — covers LiteLLM, Ollama, Perplexity, Together AI, and any other OpenAI-compatible endpoint |
byok-relay supports non-LLM APIs that return binary responses (audio, images) or accept raw audio uploads. The same BYOK model applies: your users bring their own key; byok-relay handles auth headers and binary pass-through.
| Provider | Name | Key scheme | Use cases |
|---|---|---|---|
| ElevenLabs | elevenlabs |
xi-api-key header |
Text-to-speech (TTS), speech-to-speech, voice generation |
| HuggingFace | huggingface |
Bearer token | NLP, image generation, audio models (Inference API) |
| Deepgram | deepgram |
Token scheme |
Speech-to-text (STT), text-to-speech |
Binary response handling: Audio and image responses are piped through byte-for-byte — no JSON parsing. The relay preserves Content-Type, Content-Length, and Content-Disposition headers so the client receives the raw audio/image buffer directly.
Raw audio uploads (Deepgram STT): When sending audio to /v1/listen, set Content-Type to the audio MIME type (e.g. audio/wav, audio/mpeg). The relay detects non-JSON content types and passes the raw binary body through to the provider without re-encoding.
POST /relay/elevenlabs/v1/text-to-speech/{voice_id}
x-relay-token: <your-token>
Content-Type: application/json
{ "text": "Hello from byok-relay!", "model_id": "eleven_monolingual_v1" }Response: audio/mpeg binary stream.
POST /relay/deepgram/v1/listen?model=nova-2
x-relay-token: <your-token>
Content-Type: audio/wav
<raw audio bytes>Response: JSON transcript from Deepgram.
Adding a new built-in provider is ~5 lines in src/providers.js.
| Endpoint | Description |
|---|---|
POST /users |
Register app user, get relay token |
POST /keys/:provider |
Store encrypted API key |
GET /keys |
List stored providers |
DELETE /keys/:provider |
Remove a stored key |
POST /relay |
Unified routing — model field selects provider |
GET /models |
Routing table (patterns + provider prefixes) |
POST /relay/:provider/* |
Per-provider relay (backward-compat) |
GET /health |
Health check + version |
GET /healthReturns HTTP 200 when the relay is healthy, HTTP 503 when a critical check fails.
{
"ok": true,
"version": "1.5.1",
"uptime": 3600,
"timestamp": "2026-06-11T03:00:00.000Z",
"providers": ["openai", "anthropic", "google", "groq", "openrouter", "mistral", "elevenlabs", "deepgram", "openai-compatible"],
"checks": {
"db": { "ok": true },
"config": { "ok": true, "encryption_key_set": true, "registration_gated": true }
}
}Deep / readiness probe — also pings a provider's models endpoint to verify network reachability:
GET /health?deep=1&provider=openaiAdds checks.upstream: { ok, provider, statusCode } to the response and is rate-limited more tightly than the base liveness check. Use this for post-deploy smoke tests, not per-request liveness probes because it makes an outbound network call.
Use /health as your liveness probe and /health?deep=1 as your readiness probe in K8s / docker-compose healthchecks.
POST /users
Content-Type: application/json
{ "app_id": "my-app" }→ { "token": "<relay-token>" } — store in browser localStorage
If
APP_SECRETis set, the request must includeAuthorization: Bearer <secret>:POST /users Content-Type: application/json Authorization: Bearer <APP_SECRET> { "app_id": "my-app" }Without a valid
Authorizationheader, the server returns401 Unauthorized.
POST /keys/anthropic
x-relay-token: <token>
Content-Type: application/json
{ "key": "sk-ant-..." }GET /keys
x-relay-token: <token>POST /keys/anthropic/rotate
x-relay-token: <token>
Content-Type: application/json
{ "key": "sk-ant-api03-..." }The relay validates the new key's format, pings the provider with a lightweight read-only request to confirm the key is accepted, then atomically replaces the stored key in a single DB write.
The old key is never touched if the new key fails validation or is rejected by the provider — safe to call on a live deployment.
Returns { ok: true, provider, rotated: true } if an existing key was replaced, or { ok: true, provider, rotated: false } if no prior key existed.
DELETE /keys/anthropic
x-relay-token: <token>POST /tokens/revoke
x-relay-token: <token>Immediately invalidates the token. Stored keys remain in the database but are no longer accessible. To regain access, register a new token (POST /users) and re-enter your keys.
DELETE /users
x-relay-token: <token>Permanently deletes the user account and all associated API keys. This action is irreversible.
Send a single request to POST /relay with a model field; the relay resolves
the provider automatically.
Use "provider/model-name" for an explicit route, or just the model name if it
matches a known pattern:
POST /relay
x-relay-token: <token>
Content-Type: application/json
{ "model": "anthropic/claude-3-5-haiku", "max_tokens": 256, "messages": [{"role":"user","content":"Hello"}] }POST /relay
x-relay-token: <token>
Content-Type: application/json
{ "model": "gpt-4o", "messages": [{"role":"user","content":"Hello"}] }Full streaming (SSE) is supported — pass "stream": true in the body.
Discovery: GET /models returns the full routing table plus the active model allowlist status. When unrestricted, the allowlist status is { "restricted": false, "message": "All models are permitted on this relay." }.
Body format note: the request body must match the target provider's native
API format (messages for OpenAI/Anthropic/Groq/Mistral, contents for Google).
The provider prefix is stripped from the model field before forwarding.
POST /relay/anthropic/v1/messages
x-relay-token: <token>
Content-Type: application/json
anthropic-version: 2023-06-01
{ "model": "claude-3-5-haiku-20241022", "max_tokens": 1024, "messages": [...], "stream": true }Full streaming (SSE) is supported — the response is piped directly from the provider to the browser.
POST /relay/openai-compatible/v1/chat/completions
x-relay-token: <token>
x-relay-base-url: https://openrouter.ai
Content-Type: application/json
{ "model": "...", "messages": [...] }Set ALLOWED_MODELS to a comma-separated list of model names or wildcard patterns to prevent users from requesting expensive or unsupported models. Configure the raw model value clients send, including provider prefixes for POST /relay requests that use them:
ALLOWED_MODELS=gpt-4o-mini,anthropic/claude-3-5-haiku*,google/gemini-2.0-flash*Matching is case-insensitive. * matches zero or more characters.
GET /models includes the routing table and the current allowlist status. If no allowlist is configured, the response includes:
{ "restricted": false, "message": "All models are permitted on this relay." }If an allowlist is configured, the response includes "restricted": true and "allowed_models".
If a relay request includes a model field not on the list, the relay returns:
HTTP/1.1 403 Forbidden
Content-Type: application/json
{ "error": "Model \"gpt-4o\" is not permitted on this relay.", "allowed_models": ["gpt-4o-mini", "anthropic/claude-3-5-haiku*", "google/gemini-2.0-flash*"] }# 1. Copy and fill in the env template
cp .env.example .env
# Set ENCRYPTION_SECRET (required): openssl rand -hex 32
# Set ALLOWED_ORIGINS to your frontend domain(s)
# Set APP_SECRET (strongly recommended): openssl rand -hex 32
# 2. Start the relay
docker compose up -d
# 3. Check it's healthy
docker compose ps
curl http://localhost:3000/healthSQLite data persists in the Compose named volume relay_data (mounted at /app/data inside the container).
Back up the volume contents (the SQLite file holds all encrypted API keys). Example:
docker run --rm -v relay_data:/data -v $(pwd):/out alpine sh -c \
'apk add --no-cache sqlite && sqlite3 /data/relay.db ".backup /out/relay-backup-$(date +%s).db"'Note: When you update the image, run
docker compose up --build -d— therelay_datavolume is preserved.
Pick a hosted platform based on your use case:
| Platform | Best for | Persistent storage | Cost |
|---|---|---|---|
| Railway | Production, hobby projects | ✅ Yes — volumes included | Free trial, then ~$5/mo |
| Render | Production, free tier | ✅ Yes — 1 GB disk | Free tier available |
| Vercel | Demos, prototyping | Free tier |
- Click the button above — Railway prompts for env vars
- Set
ENCRYPTION_SECRET(openssl rand -hex 32),ALLOWED_ORIGINS(your frontend domain),APP_SECRET(openssl rand -hex 32), and leaveDB_PATHas/data/relay.db - As soon as the initial deploy completes, before registering users or storing keys, open Dashboard → your service → Volumes → Add Volume and set the mount path to
/data - Redeploy, wait for
/healthto succeed, and only then use the relay — tokens and keys now survive restarts
Already used the relay without a volume? Before attaching the volume, stop writes and use Railway SSH to create a SQLite online backup of the existing
relay.dband download it. After mounting/data, restore that backup to/data/relay.dbbefore reopening the service. Keep the existingENCRYPTION_SECRETandTOKEN_HMAC_SECRET; changing either can make restored keys or tokens unusable.
- Click the button — Render reads
render.yamlfrom the repo and configures the service + 1 GB disk automatically - Override
ALLOWED_ORIGINSwith your frontend domain in the Render dashboard after deploy ENCRYPTION_SECRETandAPP_SECRETare auto-generated by Render
- Click the button — Replit imports the repo and installs dependencies automatically
- Add secrets in the Replit Secrets tab (🔒 icon in the sidebar):
ENCRYPTION_SECRET—openssl rand -hex 32ALLOWED_ORIGINS— your frontend domain (e.g.https://my-app.vercel.app)APP_SECRET—openssl rand -hex 32
- Hit Run — your relay starts at the URL shown in the Replit webview
Note: Free Repls sleep after ~5 minutes of inactivity. SQLite data persists in the Repl workspace across sleeps. For always-on production relays use Railway or Render instead.
- Click the button — Vercel clones the repo and prompts for env vars
- Set
ENCRYPTION_SECRET,ALLOWED_ORIGINS, andAPP_SECRET - Deploy — your relay is live at
https://byok-relay-<hash>.vercel.app
⚠️ Vercel limitation: Vercel serverless functions run on an ephemeral filesystem. SQLite state (registered users, stored keys) resets between cold starts. Use Vercel for demos and local testing only. For real users, deploy to Railway or Render instead.
Fastest path (dev only):
export ENCRYPTION_SECRET=$(openssl rand -hex 32) ALLOWED_ORIGINS=http://localhost:5173 && npx byok-relay⚠️ Keep the sameENCRYPTION_SECRETacross restarts. If it changes, the relay cannot decrypt previously stored keys. For anything beyond a throwaway dev run, save it in a durable.env, shell profile, or secret manager. For install options, see Setup.
Clone-and-run walkthrough:
# 1. Clone and install
git clone https://github.com/avikalpg/byok-relay.git && cd byok-relay && npm install
# 2. Configure
echo "ENCRYPTION_SECRET=$(openssl rand -hex 32)" > .env
echo "ALLOWED_ORIGINS=http://localhost:5173" >> .env # replace with your browser app's origin
# 3. Start
npm start &
i=0; until curl -fsS http://localhost:3000/health >/dev/null; do i=$((i + 1)); [ "$i" -ge 30 ] && { echo "Relay did not become ready"; exit 1; }; sleep 1; done
# 4. Register a user and get a token
TOKEN=$(curl -s -X POST http://localhost:3000/users \
-H "Content-Type: application/json" \
-d '{"app_id":"test"}' | node -e "let s=''; process.stdin.on('data', d => s += d).on('end', () => console.log(JSON.parse(s).token))")
# 5. Store your Anthropic key
curl -X POST http://localhost:3000/keys/anthropic \
-H "Content-Type: application/json" \
-H "x-relay-token: $TOKEN" \
-d '{"key":"sk-ant-YOUR-KEY-HERE"}'
# 6. Relay a request (streaming)
curl -X POST http://localhost:3000/relay/anthropic/v1/messages \
-H "Content-Type: application/json" \
-H "anthropic-version: 2023-06-01" \
-H "x-relay-token: $TOKEN" \
-d '{"model":"claude-3-5-haiku-20241022","max_tokens":256,"stream":true,"messages":[{"role":"user","content":"Hello!"}]}'Option A — npx (quickest, no install)
npx byok-relay launches a standalone relay server process — you run it alongside your existing app. It is not an embedded library; it listens on a port that your frontend calls. Set env vars in your shell before running.
export ENCRYPTION_SECRET=$(openssl rand -hex 32)
export ALLOWED_ORIGINS=https://your-app.example.com # or * for dev
npx byok-relay
⚠️ Persistence:ENCRYPTION_SECRETset viaexportis ephemeral (session only). If you restart the server without the same secret, it cannot decrypt previously stored keys and all users will need to re-register their keys. Save it to a file (e.g..env) or your shell profile for persistence. If you also customizeENCRYPTION_SALT(default:byok-relay-salt), save and keep that unchanged too — both values must match to decrypt existing keys.
Option B — global install
Same standalone server as Option A, available as a persistent command.
npm install -g byok-relay
export ENCRYPTION_SECRET=$(openssl rand -hex 32)
export ALLOWED_ORIGINS=https://your-app.example.com
byok-relay
⚠️ Persistence: Same caveat as Option A — storeENCRYPTION_SECRETsomewhere durable (e.g. a.envfile or your shell's.bashrc/.zshrc) so restarts don't invalidate existing stored keys. This applies toENCRYPTION_SALTtoo if you've customized it.
Option C — clone & run
git clone https://github.com/avikalpg/byok-relay.git
cd byok-relay
npm installcp .env.example .env
# Set ENCRYPTION_SECRET (generate: openssl rand -hex 32)
# Set ALLOWED_ORIGINS to your app's domain(s)npm start# Copy service file
sudo cp deploy/byok-relay.service /etc/systemd/system/
sudo systemctl enable --now byok-relay
# HTTPS with nginx + Let's Encrypt
sudo apt install nginx
sudo snap install --classic certbot
sudo certbot --nginx -d relay.yourdomain.com
# Add a deny block to your nginx site config to block direct DB file access:
# Inside your server {} block, add:
# location ~* \.db(-wal|-shm)?$ { deny all; return 404; }
# Then: sudo nginx -t && sudo systemctl reload nginx| Threat | Protection |
|---|---|
| API key leaked from DB backup or LFI | AES-256-GCM encryption at rest; key is never returned to clients or persisted in plaintext |
| Relay token leaked from database | HMAC-SHA256 stored token hash; raw token sent to user exactly once at registration; legacy hashes are upgraded lazily. Browser-stolen raw tokens remain usable until expiry or revocation |
| Unauthenticated registration abuse | APP_SECRET gate on POST /users when configured; rate-limited to 10 registrations/hour per IP while the limiter store is available |
SSRF via openai-compatible base URL |
URL blocklist (RFC-1918, link-local, cloud IMDS, IPv6 loopback, IPv4-mapped IPv6); HTTPS-only; DNS rebinding protection via resolved-IP validation |
| Request floods | Three-layer rate limiting: 100 req/min global, 20 AI req/min per token, 10 registrations/hour per IP. Redis-backed for serverless/multi-process deployments; limits fail open if Redis/store is unavailable |
| Unexpected expensive model usage | Optional ALLOWED_MODELS allowlist with exact names and * wildcards; rejects configured JSON relay requests whose model is outside the list |
| Path traversal beyond inference | Allowlist of permitted path prefixes per provider (/chat/completions, /completions, /embeddings, /messages, etc.) |
| Header injection into upstream requests | CRLF sanitisation on all forwarded header values |
| Hung upstream connections | 30 s AbortController hard timeout on every fetch() to AI providers |
| Token theft → permanent access | Tokens expire after 90 days (TOKEN_EXPIRY_DAYS); POST /tokens/revoke for immediate invalidation |
| WAL file exposure via nginx misconfiguration | Nginx deny rules for .db, .db-wal, and .db-shm files; DB_PATH to move DB out of web root; systemd service tightens DB file permissions |
API key storage:
scrypt(ENCRYPTION_SECRET + ENCRYPTION_SALT) → 32-byte derived key (computed once at startup)
aes-256-gcm(derived key, random 16-byte IV) → { iv, authTag, ciphertext } stored as JSON in SQLite
- Derived key cached at module scope —
scryptruns exactly once per process startup, not per request - Each key encrypted with a fresh random IV
- AES-GCM's
authTagcatches any tampering with the ciphertext ENCRYPTION_SECRETis required at startupENCRYPTION_SALTis configurable (default fallback exists for backward compat; generate your own withopenssl rand -hex 32)
Relay token storage:
HMAC-SHA256(TOKEN_HMAC_SECRET, rawToken) → tokenHash stored in SQLite
- The raw token is sent to the user exactly once (registration response) and never stored or logged
- All subsequent lookups compare
HMAC(incoming_token)against stored token hashes in SQLite - Set
TOKEN_HMAC_SECRETto use a dedicated HMAC key. Existing hashes made with the historicalENCRYPTION_SECRETfallback continue to authenticate and are upgraded lazily, provided the existingENCRYPTION_SECRETremains unchanged until every legacy-token user has authenticated and been upgraded. - Run
npm run token-migration-statuson the relay host to see conservativecurrent,legacy, and percentage counts. Existing rows begin as legacy/unconfirmed and become current after successful authentication; no user identifiers or tokens are printed. - Tokens expire after 90 days and can be revoked immediately via
POST /tokens/revoke
- Prompt content confidentiality — request bodies (prompts, conversation history) pass through the relay in plaintext on the way to AI providers. For production use with sensitive data, self-host on infrastructure you control.
- XSS in your app — the relay token lives in your app's
localStorage. An XSS vulnerability in your app can steal relay tokens. Scope tokens to IP, add CSP headers, and consider a short expiry. - Compromised
ENCRYPTION_SECRET— if your server environment is fully compromised, the encryption key is accessible. Mitigate with a cloud KMS (AWS KMS, GCP Cloud KMS) for higher assurance. - Multi-instance SQLite concurrency — SQLite handles concurrent reads well but bottlenecks on concurrent writes. For high-traffic multi-replica deployments, use a Postgres backend.
relay.byokrelay.com (managed) |
Self-hosted | |
|---|---|---|
| Setup time | Zero | ~5 min |
Control over ENCRYPTION_SECRET |
No — operator holds the key | Yes — you hold it |
| Request data flows through | Third-party infra | Your infra |
| Uptime SLA | None | Your ops |
| Good for | Prototypes, demos, development | Production, sensitive data |
For production deployments or any app with paying users: self-host. The managed relay is an easy way to evaluate byok-relay, not a production dependency.
# Required
ENCRYPTION_SECRET=$(openssl rand -hex 32) # ≥32 chars, never reuse
APP_SECRET=$(openssl rand -hex 32) # gate POST /users
TOKEN_HMAC_SECRET=$(openssl rand -hex 32) # HMAC token storage
ALLOWED_ORIGINS=https://yourdomain.com # lock down CORS
# Recommended
ENCRYPTION_SALT=$(openssl rand -hex 32) # unique per deployment; preserve with backups
REDIS_URL=redis://... # persistent rate limiting
TOKEN_EXPIRY_DAYS=30 # shorter than default 90
ALLOWED_MODELS=gpt-4o-mini,claude-haiku* # cap model access for shared/team relays
DB_PATH=/var/lib/byok-relay/relay.db # outside web root- Serve behind HTTPS (Let's Encrypt / Cloudflare)
- Restrict
ALLOWED_ORIGINSto your app's domain in production - Add nginx
denyrules for.db,.db-wal, and.db-shmfiles if DB is in the project directory - The systemd service applies
chmod 600todata/relay.db,relay.db-wal, andrelay.db-shmon every start viaExecStartPost. If deploying without systemd, runchmod 600 data/relay.db*manually after first start. - Back up SQLite safely while WAL is enabled: use SQLite's online backup mechanism, or stop the service and checkpoint the WAL before copying
relay.db(and anyrelay.db-wal/relay.db-shmfiles). The DB contains encrypted API keys; recovery requires preserving bothENCRYPTION_SECRETandENCRYPTION_SALT. - Rotate
ENCRYPTION_SECRETonly afternpm run token-migration-statusreports zero legacy or unconfirmed relay-token rows, then re-encrypt all stored API keys. API-key re-encryption alone does not migrate legacy token HMAC rows. Automated rotation tooling is not available yet; deleting users is not a safe rotation substitute because it destroys stored keys.
Report vulnerabilities through GitHub Security Advisories when available, or email the maintainer privately. Do not open public GitHub issues or post exploit details before a fix is available.
Users of relay.byokrelay.com can verify that the managed relay runs the exact public repo code:
-
Check the running commit:
curl https://relay.byokrelay.com/version # { "version": "1.0.1", "commit": "<sha>", "buildTime": "...", "repoUrl": "...", "attestationUrl": "..." } -
Download the attestation manifest for that release:
curl -L https://github.com/avikalpg/byok-relay/releases/download/v<version>/attestation.json -o attestation.json
-
Clone and verify:
git clone https://github.com/avikalpg/byok-relay byok-relay-verify cd byok-relay-verify git checkout <commit-from-step-1> node scripts/verify-attestation.js ../attestation.json # Expected: all PASS lines, exit code 0
The attestation manifest is a JSON file containing SHA-256 hashes of every file that affects relay behaviour (src/index.js, src/db.js, src/providers.js, package.json, package-lock.json). It is generated by GitHub Actions on every tagged release and attached to the release page — no trust in the relay operator required to verify it.
For higher assurance, self-host the relay on your own infrastructure using the same public code.
Two patterns, one integration:
Prosumer / individual — each user registers their own API key once. Requests use their own credits and are billed to their provider account; you spend $0 on inference. Great for developer tools, research UIs, or any product where users already have API accounts.
Team / B2B — a company admin registers the org's shared API key once. The relay token lives in your app's backend; all team members access AI through your app, which routes requests automatically. Billing, usage, and key rotation are managed inside the customer's organisation — not by you.
byok-relay handles both patterns today.
- You hold the encrypted keys — users trust your server. Mitigate with a cloud KMS-backed store for higher assurance.
- No built-in user accounts — the relay token is the only credential. Scope tokens to IP or add your own auth layer for production.
- Self-hosted — you're responsible for uptime, security updates, and backups. Or use relay.byokrelay.com and skip all of that.
- There's An AI For That — submission in review
- skills.sh — AI coding agent skill registry
- Awesome LLMOps — PR in review
- Awesome ChatGPT API — PR in review
Apache 2.0
- SECURITY.md — vulnerability reporting, incident response runbook, hardening checklist
- PRIVACY.md — privacy policy + template for operators running their own relay
Ready to integrate? → Use npx skills add avikalpg/byok-relay or point your coding agent at byokrelay.com/skill