Repository navigation
[skills-index-watchdog] Skills index is stale or degraded (fetch-failed) #136
Description
Activity
Root cause is not a network fetch failure — the whole
/docssubtree is currently serving the GitHub Pages 404, and both index workflows hit their job timeouts before the crawl finishes.Verified just now with
curl -L(unauthenticated, from outside CI):URL Result https://openamer.github.io/openamer/200 (landing page) https://openamer.github.io/openamer/docs/404 https://openamer.github.io/openamer/docs/skills404 https://openamer.github.io/openamer/docs/api/skills-index.json404 The 404 response body is 9379 bytes — the standard GitHub Pages 404 page. The file is not "failing to download"; it is not published at all. So
fetch-failedis a symptom, not the cause.Where it actually breaks
-
.github/workflows/skills-index.yml→ jobbuild-index(job limittimeout-minutes: 15). Latest run 38069872353: every step succeeds until "Build skills index" (python scripts/build_skills_index.py), which runs 16:59:15Z → 17:14:32Z and is cancelled by the 15-minute job timeout. ConsequentlyUpload index artifact→ skipped andtrigger-deploy→ skipped. No fresh artifact is ever produced. -
.github/workflows/deploy-site.yml→ jobdeploy-docs(timeout-minutes: 30). Latest run 38069872365: the "Prepare skills index" step runs 17:03:20Z → 17:33:35Z and is cancelled at 30 min. Inside that step the order is: (a)gh run downloadthe artifact from step 1 — none exists; (b)curlthe live index — 404; (c) fall back to a fresh crawlpython3 scripts/build_skills_index.py, which then consumes the entire 30-minute budget. Everything downstream (Build Docusaurus,Deploy to GitHub Pages) is skipped — which is exactly why/docscurrently has no live deployment.
This has repeated on every run since 2026-10-07 (all
cancelledat ~15m22s / ~30m15s): runs 38049049718, 37996939587, 38019110057, …Why the crawl overshoots its budget.
scripts/build_skills_index.pyperforms a full multi-source crawl on every invocation (ThreadPoolExecutor(max_workers=4)over 6+ sources — skills.sh, lobehub, clawhub, official, github, browse-sh — with per-source soft caps and httpx timeouts of 15s/30s). It also tracksrate_limited_sources; a slow or GitHub-API rate-limited upstream stretches the crawl well past the limits. There is no cache/incremental path that survives between runs, so each tick starts from zero.Fastest unblock (config only, no code change):
- Raise
build-index'stimeout-minutes(15 → 45) so the crawl can finish, thengh workflow run skills-index.ymland let it upload theskills-indexartifact. - Once that run is green,
gh workflow run deploy-site.yml -f skills_index_run_id=<that run id>.deploy-docsthen takes the artifact branch (validate_index→exit 0) and never reaches the fresh-crawl fallback, so the Pages site redeploys and/docs/**comes back. - Close this issue once
/docs/api/skills-index.jsonreturns 200 again — the 4-hourly probe will reopen it otherwise.
Longer term: give the crawl a cache (reuse the last healthy per-source snapshot) or move it into a dedicated job with a
timeout-minutes: 60, so a single slow/rate-limited upstream can no longer starve the deploy pipeline. The floors live inEXPECTED_FLOORS(~line 344 ofscripts/build_skills_index.py); the freshness probe mirrors them in.github/workflows/skills-index-freshness.yml.Refs:
scripts/build_skills_index.py,tools/skills_hub.py(source definitions),.github/workflows/skills-index.yml,.github/workflows/deploy-site.yml, output pathwebsite/static/api/skills-index.json(absent frommainright now).-
github-actions commented
on Oct 10, 2026 on Oct 10, 2026 – with GitHub ActionsContributorAuthorMore actionsProbe still failing at 2026-10-10T19:44:59Z:
fetch-failed— Could not download https://openamer.github.io/openamer/docs/api/skills-index.jsonRoot cause found — and it isn't the fetch or the network: one crawler source (
clawhub) never returns, so the build/deploy job runs into its wall-clock timeout and gets cancelled.I pulled the actual job logs for today's two runs (both on
main):build-index(skills-index.yml, run 38069872353): 16:59:28 start -> 17:14:32##[error]The operation was canceled.(15-min job timeout).deploy-docs(deploy-site.yml, run 38069872365): live index fetch already 404 -> "falling back to a fresh crawl" -> 17:03:34 start -> 17:33:33##[error]The operation was canceled.(30-min job timeout).
Both logs show the same shape:
Crawling skills.sh (sitemap)... skills.sh: 20000 unique skills (1.1s) Crawling official... official: 116 skills (0.2s) Crawling well-known... well-known: 0 skills (0.0s) Crawling github... github: 548 skills (183.9s) Crawling clawhub... <-- no result line, ever Crawling claude-marketplace... claude-marketplace: 5 skills (0.6s) Crawling lobehub... lobehub: 505 skills (0.2s) Crawling browse-sh... browse-sh: 476 skills (0.9s) ##[error]The operation was canceled.Every source prints its
N skills (t s)line within ~185 s — exceptclawhub, which never prints one. The job then sits idle for ~12 minutes (17:02:32 -> 17:14:29) until GitHub cancels it.Why
clawhubhangs, straight from the code onmain:scripts/build_skills_index.pysetsSOURCE_LIMITS = {"clawhub": 0, ...}— and the comment right there says "0 = unbounded catalog walk (max_items=0 in ClawHubSource)."ClawHubSource._load_catalog_indexgates its 12-second wall-clock budget onmax_items > 0— the comment is explicit: "Wall-clock budget is for interactive browse (max_items > 0) only. The offline index builder passes max_items=0 and must walk the full [catalog]." So on the offline path the budget is deliberatelydeadline = None-> the sequential walk runs to exhaustion with no overall cap.- I probed
clawhub.aijust now: the host is up, but/api/v1/skillstook 5.0 s and the root 2.5 s for single responses. A full ~250-request sequential walk of the 50k+ catalog at that latency needs 12-20+ minutes — more than either the 15-minbuild-indexbudget or the 30-mindeploy-docsbudget.
That closes the loop and explains why it self-reinforces: the index build overruns -> cancelled -> no
skills-indexartifact ->deploy-docsfinds the live index already 404 and falls back to a fresh crawl -> which overruns the same way -> cancelled -> nothing is ever deployed ->/docsand/docs/api/skills-index.jsonstay 404. Live status right now:/-> 200,/docs/-> 404,/docs/api/skills-index.json-> 404.Timing note: this has not been green for a while — the last successful
skills-index.ymlrun was 2026-08-26 (~6 weeks ago); every run since iscancelled.Concrete, low-risk fixes (any one unblocks it):
- Bound the offline
clawhubwalk too. Let the builder pass a deadline (or an explicit page cap) into_load_catalog_indexso a slow catalog degrades gracefully instead of hanging. Note the health gate still requiresclawhub >= 20000(here), so the cap must still surface >=20k — i.e. keep walking but under a total-seconds ceiling that fits inside the job. - Raise
timeout-minutesforbuild-index(15 -> 25) anddeploy-docs(30 -> 40) so a legitimately slow-but-complete walk can finish. Cheaper, but it re-arms itself the next time clawhub slows down. - Reuse the cached catalog. The code only caches complete walks, so every run currently pays the full price — a truncated/partial clawhub result could be cached and refreshed incrementally.
Happy to open a PR with option 1 (deadline plumbed through + a
clawhubfloor that still passes the health gate) if that's the direction you want.Refs: .github/workflows/skills-index.yml · .github/workflows/deploy-site.yml · scripts/build_skills_index.py ·
ClawHubSourcein tools/skills_hub.pyUpdate: this is now a PR — #148 bounds the offline ClawHub walk.
Fresh check today (2026-10-11), unauthenticated
curl -L:URL Result https://openamer.github.io/openamer/200 https://openamer.github.io/openamer/docs/404 https://openamer.github.io/openamer/docs/api/skills-index.json404 Still failing — and the newest
build-indexrun (38085574994, 2026-10-10 20:54Z) is yet anothercancelled, consistent with the ~15 min / ~30 min timeout analysis above.#148 implements option 1 from the earlier comment, as offered:
- New
ClawHubSource.CATALOG_WALK_OFFLINE_BUDGET_SECONDS = 840. The walk now always gets a deadline (12 s for browse, 840 s for the offline builder) instead ofdeadline = Noneon themax_items=0path. - 840 s is sized so the walk still clears the 20 k ClawHub floor at observed latency (168 pages × 200 ≈ 33 k).
- Job headroom raised:
build-index15 → 25 min,deploy-docs30 → 40 min. - A truncated walk is still not cached, so the full-catalog cache can't be poisoned.
Verified locally with real runs on the branch:
33 passed # clawhub + skills-index-health + extract-skills tests ruff: All checks passed! real-logic probe (infinite cursor, 0.1 s budget): pages fetched=6 elapsed=0.124s cached=False # was: walks to the 750-page captest_max_items_zero_respects_offline_budgetfails on currentmainand passes with the PR, so it pins the new behaviour rather than a snapshot.Once merged: dispatch
skills-index.yml, thendeploy-site.yml -f skills_index_run_id=<that run id>./docs/api/skills-index.jsonshould return 200 and this issue can close (otherwise the 4-hourly probe reopens it).🤖 Generated with OpenAmer Agent
- New
github-actions commented
on Oct 10, 2026 on Oct 10, 2026 – with GitHub ActionsContributorAuthorMore actionsProbe still failing at 2026-10-10T23:08:52Z:
fetch-failed— Could not download https://openamer.github.io/openamer/docs/api/skills-index.jsonCoordination signal on the fix path — the unblock is ready, it just needs the merge.
Fresh unauthenticated check just now (2026-10-11 ~00:52Z):
URL Result https://openamer.github.io/openamer/200 (landing page) https://openamer.github.io/openamer/docs/404 https://openamer.github.io/openamer/docs/api/skills-index.json404 And the state of the fix named in the root-cause comment above: #148 is still open and unmerged as of this check —
mergeable: true, last updated 2026-10-10T23:01Z. Neither index workflow has produced a green run since; the latestbuild-index(skills-index.yml) anddeploy-docs(deploy-site.yml) runs are stillcancelled, consistent with the ClawHub catalog walk overrunning the job timeout.So the remaining action is a single one: merge #148, then confirm one green
build-indexrun. Once/docs/api/skills-index.jsonreturns 200 the next freshness probe closes this automatically and it stops reopening — no further code change is required for this issue.Independent verification of the fix on the PR head — plus two small test-hygiene nits before merge.
Fresh check this run (~02:55Z 2026-10-11, unauthenticated
curl -L):URL Result https://openamer.github.io/openamer/200 https://openamer.github.io/openamer/docs/404 https://openamer.github.io/openamer/docs/api/skills-index.json404 Unchanged since the 00:52Z check, and no newer
build-index/deploy-docsrun has fired (latest are still 2026-10-10 20:54Z / 16:59Z, bothcancelled). So #136 stays open and the remaining act is unchanged: merge #148.To de-risk that merge I ran the PR's own test module locally against its head (
d4a7d0858):$ python -m unittest tests.tools.test_skills_hub_clawhub Ran 18 tests in 0.145s OK18/18 green, including the 6 in
TestClawHubCatalogWalkBounded. I also confirmed the change is net-new: onmainthe offline walk is stillelse None(grep forCATALOG_WALK_OFFLINE_BUDGET_SECONDSfinds nothing); on the PR head it iselse self.CATALOG_WALK_OFFLINE_BUDGET_SECONDS(=840). So the fix does exactly what the title promises.Two test-hygiene nits worth clearing while you are in there (both pass, so neither blocks — but both now describe behaviour the code no longer has):
test_max_items_zero_ignores_wall_clock_budgetpatchesCATALOG_WALK_BUDGET_SECONDS = -1and still expects 750 pages. That patch is now a no-op on the offline path (the offline path reads the offline budget), so the test passes for a different reason than its name implies — offline no longer "ignores wall clock"; it ignores only the browse budget. Suggest renaming to e.g.test_max_items_zero_ignores_the_browse_budget, or re-pointing it at the offline constant.test_max_items_zero_is_unbounded_and_caches— "unbounded" is now inaccurate;max_items=0is bounded by the 840s offline budget (it just terminates naturally in the mock here). Worth a rename so the class is self-consistent.
Both are cosmetic and I would merge as-is; flagging because the class docstring was updated for precisely this reason and these two test names were left describing the old contract.
— verified on branch
fix/skills-index-clawhub-offline-budget@d4a7d0858github-actions commented
on Oct 11, 2026 on Oct 11, 2026 – with GitHub ActionsContributorAuthorMore actionsProbe still failing at 2026-10-11T03:36:41Z:
fetch-failed— Could not download https://openamer.github.io/openamer/docs/api/skills-index.jsonFresh run, new data point — the timeout loop now reproduces on push, not just the cron, and it died at the crawl step. Also: the unblock PR is currently red.
Re-verified just now (2026-10-11 ~05:00Z, unauthenticated
curl -L):URL Result https://openamer.github.io/openamer/200 https://openamer.github.io/openamer/docs/404 https://openamer.github.io/openamer/docs/skills404 https://openamer.github.io/openamer/docs/api/skills-index.json404 What's new since the 02:58Z check: another
deploy-docsrun fired — and this one was not the 6/18 UTC cron, it was apushtomain(the darwin evolution-refresh commit768264c8, 03:05:29Z).deploy-site.ymlrun 38107279131 — started 03:05:43Z, cancelled 03:36:03Z.- Failure is localized to Step 7 "Prepare skills index (unified multi-source catalog)" (
completed / cancelled); every downstream step — build, stage, deploy —skipped. Steps 1–6 all succeeded, so this is not infra: it's still the unbounded offline ClawHub walk.
So the blast radius is wider than the cron: any push to
mainthat touches the site re-triggers a cancelled deploy, keeping/docsdark between attempts.On the unblock: #148 is still open / unmerged (head
d4a7d08) — and its checks are currently red, so it isn't a one-click merge:mergeable_state=unstable(mergeable: true), and the required gate "All required checks pass" → failure.- 6 of 8
Python testsslices fail (2, 4, 5, 6, 7, 8); only 1 and 3 pass. Review label gate→ failure (looks like the PR just needs its review label).
The head is also now behind
main(two commits landed after it). First step before merge: rebase #148 on currentmainand see whether those slices are a real regression or stale-base flakiness — then land it.Net: nothing to fix inside this issue; it stays open and correct until #148 merges, and the action item is now "rebase + fix the red slices on #148, then merge."
Found the real blocker — it's the Pages source mode, not the fetch or the network (measured 2026-10-11, ~07:04 UTC).
Reproducing the probe today, the mismatch is structural, not transient:
https://openamer.github.io/openamer/→ 200, and the bytes served are the committeddocs/index.html(the hand-written landing — the "Self-Improving AI Agent" marker matchesdocs/index.htmlonmain).https://openamer.github.io/openamer/docs/index.html→ 404https://openamer.github.io/openamer/docs/api/skills-index.json→ 404
The Pages API confirms the deployment mode:
GET /repos/openamer/openamer/pages → { "build_type": "legacy", "source": { "branch": "main", "path": "/docs" }, "status": "built", "https_enforced": true }So the site root is the committed
docs/folder — which is why the whole/docssubtree 404s: under branch mode the root would have to contain a physicaldocs/subfolder, and it doesn't. And the index is never committed —GET /repos/openamer/openamer/commits?path=docs/apireturns 0 — it only ever exists inside the Actions artifact thatdeploy-site.ymlstages as_site/docs/.Net effect: while Pages sits on
build_type: legacy, theactions/deploy-pagesstep indeploy-site.ymldoes not drive the live site, so the artifact (Docusaurus build +api/skills-index.json) never lands. The freshness probe will keep reportingfetch-failedregardless of PR #148 — the ClawHub timeout is real, but it's the second problem, not the one that makes this issue reopen.There's also a deadlock worth naming:
deploy-site.ymlfalls back to a fresh crawl whenever the live index 404s, and that crawl is what hangs → the run ends upcancelled(recentdeploy-site.ymlandskills-index.ymlruns are allconclusion: cancelled). 404 → crawl → hang → cancel → still 404.The unblock is one setting: Settings → Pages → Build and deployment → Source = "GitHub Actions", then re-run
deploy-site.yml(workflow_dispatch). That makes the_site/artifact live — Docusaurus under/docs/andskills-index.jsonat exactly the URL both the probe and the deploy's own "download live index" fallback expect. Once the artifact is live, the deploy's fallback short-circuits and stops paying the crawl cost, so the two bugs unwind together.I'll leave this open until the probe goes green and post a follow-up once the Pages source is switched.
Links: Pages settings · deploy-site.yml · skills-index.yml · PR #148
Fresh check + one gap to close before flipping the Pages source — otherwise the fix trades one 404 for another (measured 2026-10-11 ~09:10 UTC).
Re-verified this run:
GET /repos/openamer/openamer/pages→ still{ "build_type": "legacy", "source": { "branch": "main", "path": "/docs" }, "status": "built" }— unchanged since the 07:04Z comment.https://openamer.github.io/openamer/→ 200 ·/docs/→ 404 ·/docs/api/skills-index.json→ 404.- Latest
deploy-site.ymlrun is stillcancelled(03:05Z,push); PR fix(skills-hub): bound the offline ClawHub catalog walk (unblocks the Skills Index build, #136) #148 is still open and unmerged.
So the 07:04Z diagnosis holds and the unblock really is
build_type: legacy → workflow. But there is a staging gap indeploy-site.ymlthat the setting flip alone won't cover.The "Stage deployment" step only ever produces:
_site/docs/* # cp -r website/build/* (Docusaurus, baseUrl '/docs/') _site/llms.txt # cp website/build/llms.txt _site/llms-full.txtThere is no
_site/index.htmland no root-level pages. Underbuild_type: legacythe live root is the committeddocs/folder — which is exactly whyhttps://openamer.github.io/openamer/serves the hand-written landing (titleOpenAmer – Self-Improving AI Agent) and everything it links to. Flip toworkflowand the root becomes_site/, so the landing and its siblings regress to 404:currently 200 at root referenced by the landing as docs/index.html(the landing)/docs/superiority.html./superiority.htmldocs/academy.htmlacademy.htmldocs/consulting.htmlconsulting.htmldocs/skills-store.htmlskills-store.htmldocs/wiki/(28 files)wiki/Net: the flip fixes
/docs/**(the probe + the Docusaurus site) but silently 404s the marketing root unless those pages are also staged.Suggested order of operations:
- In
deploy-site.yml→ "Stage deployment", also copy the committed root pages into the artifact root, e.g.
cp docs/index.html _site/index.html
cp docs/{superiority,academy,consulting,skills-store}.html _site/
cp -r docs/wiki _site/wiki(plus any assets those pages reference) - Settings → Pages → Build and deployment → Source = "GitHub Actions", then
gh workflow run deploy-site.yml. - Verify all three:
/= 200 (landing intact),/docs/= 200,/docs/api/skills-index.json= 200 — then the freshness probe goes green and this closes.
Keeping this open until the probe is green.
Links: Pages settings · deploy-site.yml · skills-index.yml · PR #148
github-actions commented
on Oct 11, 2026 on Oct 11, 2026 – with GitHub ActionsContributorAuthorMore actionsProbe still failing at 2026-10-11T10:31:24Z:
fetch-failed— Could not download https://openamer.github.io/openamer/docs/api/skills-index.jsonFresh check + the one detail that makes this self-perpetuating — our own learning loop is re-arming the doomed deploy (measured 2026-10-11 ~13:10 UTC).
State right now (unauthenticated
curl -L):URL Result https://openamer.github.io/openamer/200 https://openamer.github.io/openamer/docs/404 https://openamer.github.io/openamer/docs/api/skills-index.json404 GET /repos/openamer/openamer/pages-> still{ "build_type": "legacy", "source": { "branch": "main", "path": "/docs" } }. PR #148 is still open withmergeable_state: unstable(6/8Python testsslices red). So both earlier diagnoses — the Pages source mode (07:04Z) and the missing root-page staging (09:07Z) — still hold and are still the merge-path blockers.What's new: the thing re-triggering the cancelled deploy is not only the 6/18 UTC cron — it is our own cron auto-commit.
deploy-site.yml'spushfilter includesskills/**, and today'schore(reports): cron auto-commitcommits addskills/auto-generated/auto-*/SKILL.md. So every learning tick that writes an auto-generated skill matches the path filter and kicks off a fresh 30-minutedeploy-docs:- run #163 — created 12:36:17Z, commit
84853cc3(skills/auto-generated/auto-internet-learning-google-to/SKILL.md) -> cancelled. - run #164 — created 12:37:24Z, commit
d65afa80(3xskills/auto-generated/.../SKILL.md) -> in_progress right now, sitting on Step 7 "Prepare skills index" (started 13:06:41Z), so it dies at the 30-min timeout ~13:36Z unless the pipeline changed.
Why Step 7 always crawls on a push: with no
skills_index_run_idit fetches the live index (deploy-site.ymlline 145) — which is 404 — then falls back topython3 scripts/build_skills_index.py(line 153), i.e. the unbounded offline ClawHub walk this issue is about. A push-triggered deploy should not be doing a full multi-source crawl, yet that is exactly its default.Independent mitigation (does not need #148 to merge) — either of these breaks the extra cancellations on its own:
- Add
paths-ignore: ['skills/auto-generated/**']to thepushfilter (or move auto-generated skills out ofskills/**) so the learning loop stops triggering site deploys at all. The auto-generated skill output does not actually need to ship the Pages site. - In Step 7, when
SKILLS_INDEX_RUN_IDis empty and the live index is unavailable, fail soft / skip instead of doing a fresh crawl (keep the full crawl only behindREBUILD_SKILLS_INDEX=true).
Either way, merging #148 is still needed for the scheduled
skills-index.ymlpath to go green; the above just stops the auto-commit side from adding a cancelled 30-min run every few hours while that lands.Refs: deploy-site.yml (on.push.paths; step 7 lines 145/153) - #148
- run #163 — created 12:36:17Z, commit
github-actions commented
on Oct 11, 2026 on Oct 11, 2026 – with GitHub ActionsContributorAuthorMore actionsProbe still failing at 2026-10-11T17:14:29Z:
fetch-failed— Could not download https://openamer.github.io/openamer/docs/api/skills-index.jsonNew since the 13:11Z check — the Pages site picked up a custom domain (
openamer.com), which changes the probe's URL surface but not the 404 (measured 2026-10-11 ~17:20 UTC).Fresh, unauthenticated checks this run:
URL Result https://openamer.github.io/openamer/301 -> https://openamer.com/(200)https://openamer.github.io/openamer/docs/301 -> https://openamer.com/docs/-> 404https://openamer.com/200 (the committed landing: <title>OpenAmer - Self-Improving AI Agent</title>)https://openamer.com/docs/404 https://openamer.com/docs/api/skills-index.json404 GET /repos/openamer/openamer/pagesnow returns:{ "cname": "openamer.com", "html_url": "https://openamer.com/", "build_type": "legacy", "source": { "branch": "main", "path": "/docs" }, "status": "built", "https_enforced": true, "https_certificate": { "state": "approved", "domains": ["openamer.com"] } }So the domain is set and the cert is approved - but
build_typeis stilllegacy, i.e. unchanged from the 07:04Z diagnosis. The github.io host is now a permanent 301 toopenamer.com, so the freshness probe's hard-coded URL is doubly wrong: it 301s toopenamer.com/docs/api/skills-index.json, which 404s. The probe usescurl -fsSL(follows redirects), so it still reportsfetch-failed- but it is no longer measuring the host you actually care about.Two concrete follow-ups that the custom-domain change adds:
-
Point the probe at the canonical host.
.github/workflows/skills-index-freshness.ymlhard-codeshttps://openamer.github.io/openamer/docs/api/skills-index.json. Switch it tohttps://openamer.com/docs/api/skills-index.jsonso the check measures the production host and a future host/redirect change can't silently mask a regression. -
The
legacy -> workflowflip is now a production-domain risk, not just a github.io one. The domain root currently is the committeddocs/folder (that's whyopenamer.com/-> 200 with the hand-written landing whileopenamer.com/docs/-> 404). The "Stage deployment" step (per the 09:07Z analysis) only produces_site/docs/*+llms.txt- there is no_site/index.html. So flipping Pages to "GitHub Actions" before the 09:07Z staging fix would 404 the marketing root onopenamer.comitself, not merely on github.io. The staged root pages (index.html,superiority.html,academy.html,consulting.html,skills-store.html,wiki/) must land in_site/first.
Net: action items are unchanged from 07:04Z/09:07Z - merge #148 (still open,
build_typestilllegacy), flip the Pages source, stage the root pages - plus one new one: retarget the freshness probe toopenamer.com. Thelegacysource mode is what keeps/docs/**404 on the new domain; the domain change alone does not touch that.-
Live re-measure at 2026-10-11 ~19:20 UTC + the actual content of the "6/8 Python slices red" blocker on #148 — it is repo-wide red, not #148's diff.
Unauthenticated
curl -Lright now:URL Result https://openamer.github.io/openamer/200 https://openamer.github.io/openamer/docs/404 https://openamer.github.io/openamer/docs/api/skills-index.json404 https://openamer.com/docs/404 https://openamer.com/docs/api/skills-index.json404 GET /repos/openamer/openamer/pages->build_type: legacy,source: main:/docs,status: building. Apages build and deploymentrun completed success at 19:07:54Z (run 38166709836) and/docs/**is still 404. Same root cause as the 07:04Z / 09:07Z checks: inlegacymode the build publishes the contents ofdocs/at the domain root (openamer.com/= the committed landing, 200), so a/docs/path cannot exist. Nothing on the merge path has changed.New this run — I opened the failing slice logs instead of re-asserting "6/8 red". On #148's head (
d4a7d08) and on a current-main run (38165823728, 18:55Z) the slices fail on pre-existing, unrelated tests — none of them touch what #148 changes:- 2/8 —
test_feishu_load_settings_ignores_extra_allow_bots(assert 'all' == 'none'),test_default_compact_banner_keeps_legacy_openamer_openamer_branding,test_kanban_db_has_no_security_findings_if_present - 4/8 —
tests/test_code_review_bot.py(3),tests/tools/test_video_generation_tool_surface_matrix.pyxai edit/extend (4) - 5/8 —
test_auth_xai_oauth_provider,test_swarm_orchestrator(2),tests/tools/test_asi_darwin.py(3) - 6/8 —
tests/scripts/test_scoreboard.py::test_every_capability_value_names_its_source - 7/8 —
test_auth_codex_quota_probe.py(3),test_autonom_watchtower,test_test_suite_watchdog.py(6) - 8/8 —
test_darwin_engine,test_swarm_home_consistency(2),test_image_generation_env(3),test_subprocess_stdin_guard,test_windows_native_support
That is ~35 distinct failures across ~14 files. The current-main run fails the same slices (2,3,4,5,7,8), so main is red and #148 merely inherits it. Conclusion: "merge #148" needs those ~35 tests green first (or a maintainer decision to relax the required check) — the ClawHub-walk fix in #148 is necessary but not sufficient on its own.
One tempting-looking shortcut worth ruling out: #157 pins
PYTHONUTF8=1in the hermetic runner and fixes the Windows canonical-suite red (507 failing locally). That is a different red from these Linux CI slices — the failures above reproduce onubuntu-latest, whereutf8_modeis already 1. #157 is a good fix, but by itself it will not clear #148's slice failures (its own CI run 38161062578 is red on the same slices).Net action list: (1) clear the ~35 pre-existing slice failures (new item this run); (2) merge #148; (3) flip Pages to
build_type: workflowonly after the 09:07Z root-page staging lands; (4) retarget the freshness probe to the canonicalopenamer.comhost. Site build lives under website/.- 2/8 —
Automated freshness probe failed.
Status:
fetch-failedDetail: Could not download https://openamer.github.io/openamer/docs/api/skills-index.json
The Skills Hub at /docs/skills depends on
/docs/api/skills-index.json.The unified index is rebuilt by
.github/workflows/skills-index.yml(cron 6/18 UTC)and
.github/workflows/deploy-site.yml(on every push affecting website/skills).If this issue keeps reopening, check the latest runs:
This issue was opened by
.github/workflows/skills-index-freshness.yml. Close it once the underlying problem is fixed; the next probe will reopen if it's still broken.