Turns curated learning resources into a Learner-Neutral Core Concept Graph, derives prerequisite structure and study assets from it, and serves them to a learner app as a playable expedition.
- AGENTS.md — engineering workflow and enforcement rules
- CONTEXT.md — project language and ambiguity resolution
- docs/adr/ — durable architectural decisions
- docs/plans/ — active plans, live TODO, and blockers
Apps:
apps/kg-worker: extraction, graph-version build, and enrichment CLIapps/learner-api: typed learner HTTP API (Hono)apps/learner-app: universal Expo learner app (web + Android)apps/admin-lab: Next.js operator inspection surface
Packages:
packages/domain-core: learner-neutral Concepts, Concept Evidence Profiles, and graph versionspackages/ports: explicit application boundariespackages/application: use-cases, projections, and orchestrationpackages/infrastructure-ingestion: structured text, HTML, and Docling parser adapterspackages/infrastructure-litellm: forced named tool-call gateway and stage descriptorspackages/infrastructure-postgres: code-first persisted schema, its generated baseline, and the one schema migratorpackages/infrastructure-storage-local: local curated-source object store adapter
cp .env.example .env
docker compose up -d postgres litellm
pnpm install
set -a; . ./.env; set +a # the shell does not auto-load .env; DB commands need DATABASE_URL
pnpm db:migrate
pnpm dev:admin # Admin Lab (Next.js)
pnpm dev:learner # Learner app web (Expo, no browser auto-open)docker compose up -d --build starts the stack. Compose brings the application schema to current
through the one-shot migrate service after PostgreSQL becomes healthy, and starts the learner API
only after that migration and LiteLLM both succeed.
Google sign-in does not work from pnpm dev:learner against the deployed API, and no setting
fixes it: localhost:8881 is cross-site with api.lrnki.globesoul.com, so the browser drops the
OAuth state cookie and every callback fails with state_mismatch
(ADR-0041). Email + password —
the path the rigs drive — is unaffected. To exercise Google, run the API yourself so that it shares
the web origin's host and scheme:
pnpm dev:api # the learner-api container, published on :8787 for this machine only
pnpm dev:learner:local # web on :8881, pointed at it (--clear: Metro inlines and caches the origin)One-time setup: register http://localhost:8787/auth/callback/google as an authorized redirect URI
on the Google client — Google exempts localhost from its https-only rule. pnpm dev:api is
docker compose watch over docker-compose.dev.yml, whose entire content is a loopback publish of
the port that container already serves: the base file gives learner-api no host port, because on
the VPS only Caddy reaches it
(ADR-0040). Edits under
apps/learner-api/src and packages/ sync into the running container and restart it; a
pnpm-lock.yaml change rebuilds the image. It runs attached — if it is not on screen it is not
syncing.
An ambient NODE_ENV in your shell breaks pnpm build. Next.js reads it directly, so a value
exported by a shell profile or left over from an earlier command sends the Admin Lab build down a
configuration path it was never meant to take, and the failure reads as a Next.js defect rather than
an environment one. Leave NODE_ENV unset and let each tool choose its own.
The caddy service is behind the public profile, so that command skips it on a development
machine — Caddy only makes sense where api.lrnki.globesoul.com resolves, and anywhere else it
retries ACME against the real VPS forever. The shared host opts in with COMPOSE_PROFILES=public in
its .env; scripts/deploy-learner-api.sh names caddy explicitly, which activates the profile on
its own, so the deploy works with or without that variable.
Run the quality checks:
pnpm checkpnpm check includes the intercepted production-web Playwright gate (pnpm e2e:web), which mocks
the API and runs deterministically. The OAuth-return case can also be rerun manually against the
deployed Pages artifact; it intercepts the deployed bundle's session read, performs no sign-in or
real API request, and touches no database:
pnpm e2e:web:deployedTwo heavier suites are opt-in and not part of pnpm check
(ADR-0038):
pnpm e2e:web:realuse # real supervisor-free API over Postgres, no generation
pnpm e2e:native:maestro # real Android APK on an emulator, deterministic loopback fixtureThe real-backend web gate needs live Postgres with at least one ready catalog enrichment; it selects one by capability, never generates, and cleans up its disposable learners on success or failure. See apps/learner-app/e2e-realuse/README.md.
The native gate drives a standalone e2e-profile APK on a booted Android emulator with Maestro. What a green run does and does not prove is owned by ADR-0038; prerequisites and setup are in apps/learner-app/e2e-native/README.md.
Persisted shape is code-first: edit the internal Drizzle schema, regenerate the sole baseline, and reset rather than add a second migration (ADR-0039). Four commands cover the whole loop:
pnpm db:generate # after editing packages/infrastructure-postgres/src/schema/ — offline
pnpm db:check # offline drift gate; already runs inside `pnpm check`
pnpm db:migrate # bring DATABASE_URL's database to current
pnpm db:reset # drop + recreate its public/drizzle schemas, then migratedb:migrate and db:reset need DATABASE_URL, which the shell does not auto-load
(set -a; . ./.env; set +a). db:generate replaces the SQL, snapshot, and journal together —
review all three, never hand-edit them, and never apply the SQL with psql.
The migrator applies the baseline to an empty database and is a no-op on a current one. Every other
state stops it before any DDL and names itself — legacy-schema, partial-schema,
stale-baseline, metadata-without-schema, or unexpected-history
(ADR-0039 holds the normative
state machine). The operator response is the same for all five: reset.
Locally that is pnpm db:reset. On the shared environment it is the cutover runbook below. A deploy
never resolves these states by itself, and no fix is ever a volume deletion — postgres_data also
holds LiteLLM's database and its virtual keys.
Runbook for the topology decided in ADR-0035 and ADR-0011.
| Surface | URL | How it deploys |
|---|---|---|
| Learner web | https://lrnki.globesoul.com |
.github/workflows/deploy-learner-web.yml on push to main |
| Learner API | https://api.lrnki.globesoul.com |
scripts/deploy-learner-api.sh (manual, run on the VPS) |
Both hostnames are stable and hardcoded in the single file that consumes each (workflow
EXPO_PUBLIC_LEARNER_API_URL, the apps/learner-api/src/app.ts CORS default,
scripts/docker/caddy/Caddyfile).
There is one shared learner environment during testing
(ADR-0036), so
pnpm --filter @lrnki/learner-app start needs no configuration. EXPO_PUBLIC_LEARNER_API_URL is
the single opt-in override for pointing the app at some other API.
Containers. Plain docker compose runs lrnki-postgres (5433), lrnki-litellm (4000),
lrnki-docling (5001), lrnki-caddy (80/443), and lrnki-learner-api, which carries the generation
supervisors and publishes no port — reach it from inside:
docker exec lrnki-learner-api node -e '…fetch("http://127.0.0.1:8787"…)'API dev loop — the public hostname has exactly one upstream, the learner-api container
(ADR-0040). Edit on the host,
run in the container:
docker compose watch learner-api # foreground; syncs src/ and packages/ into the container
docker compose logs -f learner-api # the positive signal: one restart per editA saved edit syncs and restarts the container in ~1–3s (a brief 502 during the restart is expected);
a pnpm-lock.yaml change triggers a real image rebuild instead, since a dependency change cannot be
satisfied by copying files. Watch is foreground and attached — if it is not on your screen it is
not syncing, and it dies with its terminal or SSH session. Stop it before deploying;
scripts/deploy-learner-api.sh refuses while one is attached, because a sync would overwrite the
image it just deployed.
The Caddyfile is baked into the built caddy image rather than bind-mounted (reason in
scripts/docker/caddy/Dockerfile), so a Caddy config change needs
docker compose up -d --build caddy.
Every compose command here must be detached and must run from this checkout on the host. A bare
docker compose up is attached: it takes the whole stack down when its terminal or SSH session
ends, which is how the shared environment went dark on 2026-08-05. Running compose from an agent
container that binds the workspace at a different prefix is the other half of the same rule — the
file binds set create_host_path: false and refuse such a caller by name, but watch and down
are not protected (ADR-0040,
AGENTS.md rule 23).
API deploy — from the repo checkout on the VPS (drives the local Docker daemon):
scripts/deploy-learner-api.sh # git pull → build → migrate (verified) → up learner-api caddy → probe container, then public /healthThe learner-api container reads DATABASE_URL and LITELLM_BASE_URL from compose;
LITELLM_API_KEY, BETTER_AUTH_SECRET, BETTER_AUTH_URL, GOOGLE_CLIENT_ID, and
GOOGLE_CLIENT_SECRET come from the repo-root .env. The deploy brings the schema to current
through the one-shot migrate service and aborts before touching the API if that container exits
nonzero, so a healthy old API can never report a successful deploy over a failed migration. A
migration that reports reset-required is never resolved by the deploy — it waits for the explicit
reset runbook. Learner sessions are Better Auth rows in session
(ADR-0041) and survive
restarts; BETTER_AUTH_SECRET signs their cookies, so rotating it signs every learner out.
BETTER_AUTH_URL must be the API's public origin. Better Auth derives both the Google redirect URI
it advertises (${BETTER_AUTH_URL}/auth/callback/google) and, from that URL's scheme, whether
session cookies carry Secure — so the .env.example dev default left on this host deploys an API
that health-checks green and serves the whole credential path while Google rejects the callback and
every session cookie ships without Secure over HTTPS. Nothing errors, because a wrong base URL
still resolves. The deploy now asserts the shipped value against the origin it serves and fails
loudly on a mismatch; curl -sSI the Set-Cookie from a sign-up if you need to confirm by hand.
The deploy does not reload LiteLLM. It rebuilds migrate, learner-api, and caddy only.
litellm/config.yaml is a read-only bind read once at process start, and store_model_in_db is
unset, so the file is authoritative only at that moment — a commit that repoints a
model_group_alias leaves the running router serving the previous model, with no error anywhere,
because the alias still resolves. That is a silent stale deploy, and it happened: the
kg-independent-judge → deepseek-v4-flash-0731 swap of 2026-08-07 did not take effect on the
shared host until 2026-08-08, so every judge call in between ran the model it replaced. After a
deploy whose range touches litellm/config.yaml, reload it and confirm the new deployment is
actually served:
docker compose up -d --force-recreate --no-deps litellm # never `down -v`: it holds LiteLLM's keys
curl -s -H "Authorization: Bearer $LITELLM_API_KEY" http://127.0.0.1:4000/models | grep -c <new-model>A 200 from /models is not evidence the config is current — check for the deployment by name. The
router's own answer is in LiteLLM_SpendLogs.model, which records which deployment actually served.
Shared schema cutover — the only response to a reset-required deploy, and deliberately manual.
It discards the application data in database lrnki (greenfield: no backup or data migration is
an acceptance dependency) and preserves everything else. From the repo checkout on the VPS:
docker compose stop learner-api # stop the writers
docker compose exec -T postgres \
psql -U lrnki -d lrnki -X -v ON_ERROR_STOP=1 < scripts/reset-app-schema.sql
scripts/deploy-learner-api.sh # migrate applies 0000 once, then the APINever pipe that psql into anything — a pipeline reports the last command's status, which would
hide the guard. scripts/reset-app-schema.sql aborts on any database other than lrnki/lrnki_test
(exit 3) and drops only the public and drizzle schemas, so the litellm database sharing the
postgres_data volume survives. Never docker compose down -v: that destroys LiteLLM's virtual
keys, and a dead sk-… then fails generation with 401 while LITELLM_MASTER_KEY still works.
Then confirm the cutover, including that the separate LiteLLM database survived:
docker compose exec -T postgres psql -U lrnki -d lrnki -X -Atqc \
'select count(*) from drizzle.__drizzle_migrations;' # exactly 1
curl -fsS https://api.lrnki.globesoul.com/health
curl -fsS -H "Authorization: Bearer ${LITELLM_API_KEY}" http://127.0.0.1:4000/models >/dev/nullplus one authenticated learner read/write against the API.
Verifying a rebuild — learner-api can look rebuilt and not be, in two independent ways:
- Never pipe the build.
docker compose up -d --build --no-deps learner-api | tailreports tail's exit code, so a build that died onno space left on devicestill looks like exit 0. Reclaim withdocker builder prune -f. - The container's
.Createdis not the image's. Prove the recreate by comparingdocker inspect lrnki-learner-api --format '{{.Image}}'againstdocker image inspect lrnki-learner-api:latest --format '{{.Id}}'. A stale container serves the previous behaviour while every probe passes.
Telling a dead LiteLLM key from an upstream problem — LiteLLM's virtual keys live in its own
database inside the shared postgres_data volume, so a re-initialised volume leaves the sk-… in
.env pointing at a key that no longer exists. The two failures are distinguishable:
| Symptom | Cause | Remedy |
|---|---|---|
Generation 401 while LITELLM_MASTER_KEY still works |
Dead virtual key — the master key is validated from config rather than the key table, and that asymmetry is the tell | Mint one via POST /key/generate with the master key, write it to .env, and recreate the container: container env is fixed at creation, so docker restart will not pick it up |
429 "No deployments available" |
Upstream provider rate limit, not a key problem | Wait and retry; read the run against the throttling signatures in .agents/skills/real-use-quality-evaluation/SKILL.md |
Keep any .env backup outside the repo: .gitignore covers .env but not .env.bak-*.
Reading docker logs lrnki-litellm — tell host-run tools from the container by source IP, not
message text: the container is 172.18.0.5, anything on the host (admin-lab, kg-worker) is the
gateway 172.18.0.1. A host process reads .env once at start, so a session started before a key
repair keeps presenting the dead key while the container has already picked up the new one — and the
two are identical in the message text.
Web deploy — automatic on push to main. lrnki.globesoul.com is attached as the Pages custom
domain, so the default sergkhl.github.io/lrnki/ URL 301s to it, and Enforce HTTPS is on.
Mobile builds (Android) — .github/workflows/build-learner-android.yml
(workflow_dispatch) runs scripts/build-learner-android.sh on a GitHub runner (eas build --local, authenticated by the EXPO_TOKEN repo secret) and uploads the APK as a workflow
artifact; download and sideload it. Profiles come from apps/learner-app/eas.json: preview
(default; standalone APK against the live API) and development (dev client for expo start).
Local fallback on a machine with Java 17 + the Android SDK (EXPO_TOKEN is read from the
repo-root .env, else the environment):
pnpm build:android # preview profile
pnpm build:android:dev # development profileNative dev loop — build, install, and launch a development build on a connected device (else the emulator/simulator) and start Metro with the dev client. Both target the local API, so start it first:
pnpm dev:api # the learner-api they point at
pnpm dev:android # needs Java 17 + Android SDK
pnpm dev:ios # needs macOS + Xcodeexpo run:* generates apps/learner-app/android/ and ios/ in-tree (gitignored); no EXPO_TOKEN
needed. These and dev:learner:local all go through scripts/dev-learner-app.sh, which points the
app at whatever BETTER_AUTH_URL names in .env — deliberately the same value the API signs and
advertises to Google, since a build pointed anywhere else cannot complete the sign-in leg
(ADR-0041). Android also gets
an adb reverse for that port, so localhost inside the device means this machine; that works on a
USB-attached physical device too, where the 10.0.2.2 emulator alias does not. With no device
attached the script boots an AVD (ANDROID_AVD, else the first one) and blocks until it answers as
booted — expo run:android installs as soon as a serial appears, and treats every emulator as ready
whether or not it is. Because a quick-boot guest can still drop system_server during the build
that follows, the script re-runs once when the install fails with Can't find service: package, and
never for any other reason. Debug builds permit
cleartext HTTP, so no config change is needed. Export EXPO_PUBLIC_LEARNER_API_URL to override —
that is how you point a native build at the deployed API, which nothing binds locally
(ADR-0040).
EAS iOS builds (distributable artifacts) — future work.
Course planning, OCR, multimodal interpretation, and automatic ontology import.