Skip to content

[skills-index-watchdog] Skills index is stale or degraded (fetch-failed) #136

Description

@github-actions

Automated freshness probe failed.

Status: fetch-failed
Detail: Could not download https://openamer.github.io/openamer/docs/api/skills-index.json

The Skills Hub at /docs/skills depends on /docs/api/skills-index.json.
The unified index is rebuilt by .github/workflows/skills-index.yml (cron 6/18 UTC)
and .github/workflows/deploy-site.yml (on every push affecting website/skills).
If this issue keeps reopening, check the latest runs:

This issue was opened by .github/workflows/skills-index-freshness.yml. Close it once the underlying problem is fixed; the next probe will reopen if it's still broken.

Activity

  1. openamer commented on Oct 10, 2026

    @openamer
    Owner

    Root cause is not a network fetch failure — the whole /docs subtree is currently serving the GitHub Pages 404, and both index workflows hit their job timeouts before the crawl finishes.

    Verified just now with curl -L (unauthenticated, from outside CI):

    URL Result
    https://openamer.github.io/openamer/ 200 (landing page)
    https://openamer.github.io/openamer/docs/ 404
    https://openamer.github.io/openamer/docs/skills 404
    https://openamer.github.io/openamer/docs/api/skills-index.json 404

    The 404 response body is 9379 bytes — the standard GitHub Pages 404 page. The file is not "failing to download"; it is not published at all. So fetch-failed is a symptom, not the cause.

    Where it actually breaks

    1. .github/workflows/skills-index.yml → job build-index (job limit timeout-minutes: 15). Latest run 38069872353: every step succeeds until "Build skills index" (python scripts/build_skills_index.py), which runs 16:59:15Z → 17:14:32Z and is cancelled by the 15-minute job timeout. Consequently Upload index artifact → skipped and trigger-deploy → skipped. No fresh artifact is ever produced.

    2. .github/workflows/deploy-site.yml → job deploy-docs (timeout-minutes: 30). Latest run 38069872365: the "Prepare skills index" step runs 17:03:20Z → 17:33:35Z and is cancelled at 30 min. Inside that step the order is: (a) gh run download the artifact from step 1 — none exists; (b) curl the live index — 404; (c) fall back to a fresh crawl python3 scripts/build_skills_index.py, which then consumes the entire 30-minute budget. Everything downstream (Build Docusaurus, Deploy to GitHub Pages) is skipped — which is exactly why /docs currently has no live deployment.

    This has repeated on every run since 2026-10-07 (all cancelled at ~15m22s / ~30m15s): runs 38049049718, 37996939587, 38019110057, …

    Why the crawl overshoots its budget. scripts/build_skills_index.py performs a full multi-source crawl on every invocation (ThreadPoolExecutor(max_workers=4) over 6+ sources — skills.sh, lobehub, clawhub, official, github, browse-sh — with per-source soft caps and httpx timeouts of 15s/30s). It also tracks rate_limited_sources; a slow or GitHub-API rate-limited upstream stretches the crawl well past the limits. There is no cache/incremental path that survives between runs, so each tick starts from zero.

    Fastest unblock (config only, no code change):

    1. Raise build-index's timeout-minutes (15 → 45) so the crawl can finish, then gh workflow run skills-index.yml and let it upload the skills-index artifact.
    2. Once that run is green, gh workflow run deploy-site.yml -f skills_index_run_id=<that run id>. deploy-docs then takes the artifact branch (validate_index → exit 0) and never reaches the fresh-crawl fallback, so the Pages site redeploys and /docs/** comes back.
    3. Close this issue once /docs/api/skills-index.json returns 200 again — the 4-hourly probe will reopen it otherwise.

    Longer term: give the crawl a cache (reuse the last healthy per-source snapshot) or move it into a dedicated job with a timeout-minutes: 60, so a single slow/rate-limited upstream can no longer starve the deploy pipeline. The floors live in EXPECTED_FLOORS (~line 344 of scripts/build_skills_index.py); the freshness probe mirrors them in .github/workflows/skills-index-freshness.yml.

    Refs: scripts/build_skills_index.py, tools/skills_hub.py (source definitions), .github/workflows/skills-index.yml, .github/workflows/deploy-site.yml, output path website/static/api/skills-index.json (absent from main right now).

  2. github-actions commented on Oct 10, 2026

    @github-actions
    ContributorAuthor

    Probe still failing at 2026-10-10T19:44:59Z: fetch-failed — Could not download https://openamer.github.io/openamer/docs/api/skills-index.json

  3. openamer commented on Oct 10, 2026

    @openamer
    Owner

    Root cause found — and it isn't the fetch or the network: one crawler source (clawhub) never returns, so the build/deploy job runs into its wall-clock timeout and gets cancelled.

    I pulled the actual job logs for today's two runs (both on main):

    • build-index (skills-index.yml, run 38069872353): 16:59:28 start -> 17:14:32 ##[error]The operation was canceled. (15-min job timeout).
    • deploy-docs (deploy-site.yml, run 38069872365): live index fetch already 404 -> "falling back to a fresh crawl" -> 17:03:34 start -> 17:33:33 ##[error]The operation was canceled. (30-min job timeout).

    Both logs show the same shape:

    Crawling skills.sh (sitemap)...   skills.sh: 20000 unique skills (1.1s)
    Crawling official...              official: 116 skills (0.2s)
    Crawling well-known...            well-known: 0 skills (0.0s)
    Crawling github...                github: 548 skills (183.9s)
    Crawling clawhub...               <-- no result line, ever
    Crawling claude-marketplace...    claude-marketplace: 5 skills (0.6s)
    Crawling lobehub...               lobehub: 505 skills (0.2s)
    Crawling browse-sh...             browse-sh: 476 skills (0.9s)
    ##[error]The operation was canceled.
    

    Every source prints its N skills (t s) line within ~185 s — except clawhub, which never prints one. The job then sits idle for ~12 minutes (17:02:32 -> 17:14:29) until GitHub cancels it.

    Why clawhub hangs, straight from the code on main:

    • scripts/build_skills_index.py sets SOURCE_LIMITS = {"clawhub": 0, ...} — and the comment right there says "0 = unbounded catalog walk (max_items=0 in ClawHubSource)."
    • ClawHubSource._load_catalog_index gates its 12-second wall-clock budget on max_items > 0 — the comment is explicit: "Wall-clock budget is for interactive browse (max_items > 0) only. The offline index builder passes max_items=0 and must walk the full [catalog]." So on the offline path the budget is deliberately deadline = None -> the sequential walk runs to exhaustion with no overall cap.
    • I probed clawhub.ai just now: the host is up, but /api/v1/skills took 5.0 s and the root 2.5 s for single responses. A full ~250-request sequential walk of the 50k+ catalog at that latency needs 12-20+ minutes — more than either the 15-min build-index budget or the 30-min deploy-docs budget.

    That closes the loop and explains why it self-reinforces: the index build overruns -> cancelled -> no skills-index artifact -> deploy-docs finds the live index already 404 and falls back to a fresh crawl -> which overruns the same way -> cancelled -> nothing is ever deployed -> /docs and /docs/api/skills-index.json stay 404. Live status right now: / -> 200, /docs/ -> 404, /docs/api/skills-index.json -> 404.

    Timing note: this has not been green for a while — the last successful skills-index.yml run was 2026-08-26 (~6 weeks ago); every run since is cancelled.

    Concrete, low-risk fixes (any one unblocks it):

    1. Bound the offline clawhub walk too. Let the builder pass a deadline (or an explicit page cap) into _load_catalog_index so a slow catalog degrades gracefully instead of hanging. Note the health gate still requires clawhub >= 20000 (here), so the cap must still surface >=20k — i.e. keep walking but under a total-seconds ceiling that fits inside the job.
    2. Raise timeout-minutes for build-index (15 -> 25) and deploy-docs (30 -> 40) so a legitimately slow-but-complete walk can finish. Cheaper, but it re-arms itself the next time clawhub slows down.
    3. Reuse the cached catalog. The code only caches complete walks, so every run currently pays the full price — a truncated/partial clawhub result could be cached and refreshed incrementally.

    Happy to open a PR with option 1 (deadline plumbed through + a clawhub floor that still passes the health gate) if that's the direction you want.

    Refs: .github/workflows/skills-index.yml · .github/workflows/deploy-site.yml · scripts/build_skills_index.py · ClawHubSource in tools/skills_hub.py

  4. openamer commented on Oct 10, 2026

    @openamer
    Owner

    Update: this is now a PR — #148 bounds the offline ClawHub walk.

    Fresh check today (2026-10-11), unauthenticated curl -L:

    URL Result
    https://openamer.github.io/openamer/ 200
    https://openamer.github.io/openamer/docs/ 404
    https://openamer.github.io/openamer/docs/api/skills-index.json 404

    Still failing — and the newest build-index run (38085574994, 2026-10-10 20:54Z) is yet another cancelled, consistent with the ~15 min / ~30 min timeout analysis above.

    #148 implements option 1 from the earlier comment, as offered:

    • New ClawHubSource.CATALOG_WALK_OFFLINE_BUDGET_SECONDS = 840. The walk now always gets a deadline (12 s for browse, 840 s for the offline builder) instead of deadline = None on the max_items=0 path.
    • 840 s is sized so the walk still clears the 20 k ClawHub floor at observed latency (168 pages × 200 ≈ 33 k).
    • Job headroom raised: build-index 15 → 25 min, deploy-docs 30 → 40 min.
    • A truncated walk is still not cached, so the full-catalog cache can't be poisoned.

    Verified locally with real runs on the branch:

    33 passed   # clawhub + skills-index-health + extract-skills tests
    ruff: All checks passed!
    real-logic probe (infinite cursor, 0.1 s budget):
        pages fetched=6  elapsed=0.124s  cached=False   # was: walks to the 750-page cap
    

    test_max_items_zero_respects_offline_budget fails on current main and passes with the PR, so it pins the new behaviour rather than a snapshot.

    Once merged: dispatch skills-index.yml, then deploy-site.yml -f skills_index_run_id=<that run id>. /docs/api/skills-index.json should return 200 and this issue can close (otherwise the 4-hourly probe reopens it).

    🤖 Generated with OpenAmer Agent

  5. github-actions commented on Oct 10, 2026

    @github-actions
    ContributorAuthor

    Probe still failing at 2026-10-10T23:08:52Z: fetch-failed — Could not download https://openamer.github.io/openamer/docs/api/skills-index.json

  6. openamer commented on Oct 11, 2026

    @openamer
    Owner

    Coordination signal on the fix path — the unblock is ready, it just needs the merge.

    Fresh unauthenticated check just now (2026-10-11 ~00:52Z):

    URL Result
    https://openamer.github.io/openamer/ 200 (landing page)
    https://openamer.github.io/openamer/docs/ 404
    https://openamer.github.io/openamer/docs/api/skills-index.json 404

    And the state of the fix named in the root-cause comment above: #148 is still open and unmerged as of this check — mergeable: true, last updated 2026-10-10T23:01Z. Neither index workflow has produced a green run since; the latest build-index (skills-index.yml) and deploy-docs (deploy-site.yml) runs are still cancelled, consistent with the ClawHub catalog walk overrunning the job timeout.

    So the remaining action is a single one: merge #148, then confirm one green build-index run. Once /docs/api/skills-index.json returns 200 the next freshness probe closes this automatically and it stops reopening — no further code change is required for this issue.

  7. openamer commented on Oct 11, 2026

    @openamer
    Owner

    Independent verification of the fix on the PR head — plus two small test-hygiene nits before merge.

    Fresh check this run (~02:55Z 2026-10-11, unauthenticated curl -L):

    URL Result
    https://openamer.github.io/openamer/ 200
    https://openamer.github.io/openamer/docs/ 404
    https://openamer.github.io/openamer/docs/api/skills-index.json 404

    Unchanged since the 00:52Z check, and no newer build-index / deploy-docs run has fired (latest are still 2026-10-10 20:54Z / 16:59Z, both cancelled). So #136 stays open and the remaining act is unchanged: merge #148.

    To de-risk that merge I ran the PR's own test module locally against its head (d4a7d0858):

    $ python -m unittest tests.tools.test_skills_hub_clawhub
    Ran 18 tests in 0.145s
    OK
    

    18/18 green, including the 6 in TestClawHubCatalogWalkBounded. I also confirmed the change is net-new: on main the offline walk is still else None (grep for CATALOG_WALK_OFFLINE_BUDGET_SECONDS finds nothing); on the PR head it is else self.CATALOG_WALK_OFFLINE_BUDGET_SECONDS (=840). So the fix does exactly what the title promises.

    Two test-hygiene nits worth clearing while you are in there (both pass, so neither blocks — but both now describe behaviour the code no longer has):

    1. test_max_items_zero_ignores_wall_clock_budget patches CATALOG_WALK_BUDGET_SECONDS = -1 and still expects 750 pages. That patch is now a no-op on the offline path (the offline path reads the offline budget), so the test passes for a different reason than its name implies — offline no longer "ignores wall clock"; it ignores only the browse budget. Suggest renaming to e.g. test_max_items_zero_ignores_the_browse_budget, or re-pointing it at the offline constant.
    2. test_max_items_zero_is_unbounded_and_caches — "unbounded" is now inaccurate; max_items=0 is bounded by the 840s offline budget (it just terminates naturally in the mock here). Worth a rename so the class is self-consistent.

    Both are cosmetic and I would merge as-is; flagging because the class docstring was updated for precisely this reason and these two test names were left describing the old contract.

    — verified on branch fix/skills-index-clawhub-offline-budget @ d4a7d0858

  8. github-actions commented on Oct 11, 2026

    @github-actions
    ContributorAuthor

    Probe still failing at 2026-10-11T03:36:41Z: fetch-failed — Could not download https://openamer.github.io/openamer/docs/api/skills-index.json

  9. openamer commented on Oct 11, 2026

    @openamer
    Owner

    Fresh run, new data point — the timeout loop now reproduces on push, not just the cron, and it died at the crawl step. Also: the unblock PR is currently red.

    Re-verified just now (2026-10-11 ~05:00Z, unauthenticated curl -L):

    URL Result
    https://openamer.github.io/openamer/ 200
    https://openamer.github.io/openamer/docs/ 404
    https://openamer.github.io/openamer/docs/skills 404
    https://openamer.github.io/openamer/docs/api/skills-index.json 404

    What's new since the 02:58Z check: another deploy-docs run fired — and this one was not the 6/18 UTC cron, it was a push to main (the darwin evolution-refresh commit 768264c8, 03:05:29Z).

    • deploy-site.yml run 38107279131 — started 03:05:43Z, cancelled 03:36:03Z.
    • Failure is localized to Step 7 "Prepare skills index (unified multi-source catalog)" (completed / cancelled); every downstream step — build, stage, deploy — skipped. Steps 1–6 all succeeded, so this is not infra: it's still the unbounded offline ClawHub walk.

    So the blast radius is wider than the cron: any push to main that touches the site re-triggers a cancelled deploy, keeping /docs dark between attempts.

    On the unblock: #148 is still open / unmerged (head d4a7d08) — and its checks are currently red, so it isn't a one-click merge:

    • mergeable_state = unstable (mergeable: true), and the required gate "All required checks pass" → failure.
    • 6 of 8 Python tests slices fail (2, 4, 5, 6, 7, 8); only 1 and 3 pass.
    • Review label gate → failure (looks like the PR just needs its review label).

    The head is also now behind main (two commits landed after it). First step before merge: rebase #148 on current main and see whether those slices are a real regression or stale-base flakiness — then land it.

    Net: nothing to fix inside this issue; it stays open and correct until #148 merges, and the action item is now "rebase + fix the red slices on #148, then merge."

  10. openamer commented on Oct 11, 2026

    @openamer
    Owner

    Found the real blocker — it's the Pages source mode, not the fetch or the network (measured 2026-10-11, ~07:04 UTC).

    Reproducing the probe today, the mismatch is structural, not transient:

    • https://openamer.github.io/openamer/ → 200, and the bytes served are the committed docs/index.html (the hand-written landing — the "Self-Improving AI Agent" marker matches docs/index.html on main).
    • https://openamer.github.io/openamer/docs/index.html → 404
    • https://openamer.github.io/openamer/docs/api/skills-index.json → 404

    The Pages API confirms the deployment mode:

    GET /repos/openamer/openamer/pages
    → { "build_type": "legacy", "source": { "branch": "main", "path": "/docs" },
        "status": "built", "https_enforced": true }
    

    So the site root is the committed docs/ folder — which is why the whole /docs subtree 404s: under branch mode the root would have to contain a physical docs/ subfolder, and it doesn't. And the index is never committed — GET /repos/openamer/openamer/commits?path=docs/api returns 0 — it only ever exists inside the Actions artifact that deploy-site.yml stages as _site/docs/.

    Net effect: while Pages sits on build_type: legacy, the actions/deploy-pages step in deploy-site.yml does not drive the live site, so the artifact (Docusaurus build + api/skills-index.json) never lands. The freshness probe will keep reporting fetch-failed regardless of PR #148 — the ClawHub timeout is real, but it's the second problem, not the one that makes this issue reopen.

    There's also a deadlock worth naming: deploy-site.yml falls back to a fresh crawl whenever the live index 404s, and that crawl is what hangs → the run ends up cancelled (recent deploy-site.yml and skills-index.yml runs are all conclusion: cancelled). 404 → crawl → hang → cancel → still 404.

    The unblock is one setting: Settings → Pages → Build and deployment → Source = "GitHub Actions", then re-run deploy-site.yml (workflow_dispatch). That makes the _site/ artifact live — Docusaurus under /docs/ and skills-index.json at exactly the URL both the probe and the deploy's own "download live index" fallback expect. Once the artifact is live, the deploy's fallback short-circuits and stops paying the crawl cost, so the two bugs unwind together.

    I'll leave this open until the probe goes green and post a follow-up once the Pages source is switched.

    Links: Pages settings · deploy-site.yml · skills-index.yml · PR #148

  11. openamer commented on Oct 11, 2026

    @openamer
    Owner

    Fresh check + one gap to close before flipping the Pages source — otherwise the fix trades one 404 for another (measured 2026-10-11 ~09:10 UTC).

    Re-verified this run:

    So the 07:04Z diagnosis holds and the unblock really is build_type: legacy → workflow. But there is a staging gap in deploy-site.yml that the setting flip alone won't cover.

    The "Stage deployment" step only ever produces:

    _site/docs/*        # cp -r website/build/*   (Docusaurus, baseUrl '/docs/')
    _site/llms.txt      # cp website/build/llms.txt
    _site/llms-full.txt
    

    There is no _site/index.html and no root-level pages. Under build_type: legacy the live root is the committed docs/ folder — which is exactly why https://openamer.github.io/openamer/ serves the hand-written landing (title OpenAmer – Self-Improving AI Agent) and everything it links to. Flip to workflow and the root becomes _site/, so the landing and its siblings regress to 404:

    currently 200 at root referenced by the landing as
    docs/index.html (the landing) /
    docs/superiority.html ./superiority.html
    docs/academy.html academy.html
    docs/consulting.html consulting.html
    docs/skills-store.html skills-store.html
    docs/wiki/ (28 files) wiki/

    Net: the flip fixes /docs/** (the probe + the Docusaurus site) but silently 404s the marketing root unless those pages are also staged.

    Suggested order of operations:

    1. In deploy-site.yml → "Stage deployment", also copy the committed root pages into the artifact root, e.g.
      cp docs/index.html _site/index.html
      cp docs/{superiority,academy,consulting,skills-store}.html _site/
      cp -r docs/wiki _site/wiki (plus any assets those pages reference)
    2. Settings → Pages → Build and deployment → Source = "GitHub Actions", then gh workflow run deploy-site.yml.
    3. Verify all three: / = 200 (landing intact), /docs/ = 200, /docs/api/skills-index.json = 200 — then the freshness probe goes green and this closes.

    Keeping this open until the probe is green.

    Links: Pages settings · deploy-site.yml · skills-index.yml · PR #148

  12. github-actions commented on Oct 11, 2026

    @github-actions
    ContributorAuthor

    Probe still failing at 2026-10-11T10:31:24Z: fetch-failed — Could not download https://openamer.github.io/openamer/docs/api/skills-index.json

  13. openamer commented on Oct 11, 2026

    @openamer
    Owner

    Fresh check + the one detail that makes this self-perpetuating — our own learning loop is re-arming the doomed deploy (measured 2026-10-11 ~13:10 UTC).

    State right now (unauthenticated curl -L):

    URL Result
    https://openamer.github.io/openamer/ 200
    https://openamer.github.io/openamer/docs/ 404
    https://openamer.github.io/openamer/docs/api/skills-index.json 404

    GET /repos/openamer/openamer/pages -> still { "build_type": "legacy", "source": { "branch": "main", "path": "/docs" } }. PR #148 is still open with mergeable_state: unstable (6/8 Python tests slices red). So both earlier diagnoses — the Pages source mode (07:04Z) and the missing root-page staging (09:07Z) — still hold and are still the merge-path blockers.

    What's new: the thing re-triggering the cancelled deploy is not only the 6/18 UTC cron — it is our own cron auto-commit.

    deploy-site.yml's push filter includes skills/**, and today's chore(reports): cron auto-commit commits add skills/auto-generated/auto-*/SKILL.md. So every learning tick that writes an auto-generated skill matches the path filter and kicks off a fresh 30-minute deploy-docs:

    • run #163 — created 12:36:17Z, commit 84853cc3 (skills/auto-generated/auto-internet-learning-google-to/SKILL.md) -> cancelled.
    • run #164 — created 12:37:24Z, commit d65afa80 (3x skills/auto-generated/.../SKILL.md) -> in_progress right now, sitting on Step 7 "Prepare skills index" (started 13:06:41Z), so it dies at the 30-min timeout ~13:36Z unless the pipeline changed.

    Why Step 7 always crawls on a push: with no skills_index_run_id it fetches the live index (deploy-site.yml line 145) — which is 404 — then falls back to python3 scripts/build_skills_index.py (line 153), i.e. the unbounded offline ClawHub walk this issue is about. A push-triggered deploy should not be doing a full multi-source crawl, yet that is exactly its default.

    Independent mitigation (does not need #148 to merge) — either of these breaks the extra cancellations on its own:

    1. Add paths-ignore: ['skills/auto-generated/**'] to the push filter (or move auto-generated skills out of skills/**) so the learning loop stops triggering site deploys at all. The auto-generated skill output does not actually need to ship the Pages site.
    2. In Step 7, when SKILLS_INDEX_RUN_ID is empty and the live index is unavailable, fail soft / skip instead of doing a fresh crawl (keep the full crawl only behind REBUILD_SKILLS_INDEX=true).

    Either way, merging #148 is still needed for the scheduled skills-index.yml path to go green; the above just stops the auto-commit side from adding a cancelled 30-min run every few hours while that lands.

    Refs: deploy-site.yml (on.push.paths; step 7 lines 145/153) - #148

  14. github-actions commented on Oct 11, 2026

    @github-actions
    ContributorAuthor

    Probe still failing at 2026-10-11T17:14:29Z: fetch-failed — Could not download https://openamer.github.io/openamer/docs/api/skills-index.json

  15. openamer commented on Oct 11, 2026

    @openamer
    Owner

    New since the 13:11Z check — the Pages site picked up a custom domain (openamer.com), which changes the probe's URL surface but not the 404 (measured 2026-10-11 ~17:20 UTC).

    Fresh, unauthenticated checks this run:

    URL Result
    https://openamer.github.io/openamer/ 301 -> https://openamer.com/ (200)
    https://openamer.github.io/openamer/docs/ 301 -> https://openamer.com/docs/ -> 404
    https://openamer.com/ 200 (the committed landing: <title>OpenAmer - Self-Improving AI Agent</title>)
    https://openamer.com/docs/ 404
    https://openamer.com/docs/api/skills-index.json 404

    GET /repos/openamer/openamer/pages now returns:

    { "cname": "openamer.com", "html_url": "https://openamer.com/",
      "build_type": "legacy", "source": { "branch": "main", "path": "/docs" },
      "status": "built", "https_enforced": true,
      "https_certificate": { "state": "approved", "domains": ["openamer.com"] } }

    So the domain is set and the cert is approved - but build_type is still legacy, i.e. unchanged from the 07:04Z diagnosis. The github.io host is now a permanent 301 to openamer.com, so the freshness probe's hard-coded URL is doubly wrong: it 301s to openamer.com/docs/api/skills-index.json, which 404s. The probe uses curl -fsSL (follows redirects), so it still reports fetch-failed - but it is no longer measuring the host you actually care about.

    Two concrete follow-ups that the custom-domain change adds:

    1. Point the probe at the canonical host. .github/workflows/skills-index-freshness.yml hard-codes https://openamer.github.io/openamer/docs/api/skills-index.json. Switch it to https://openamer.com/docs/api/skills-index.json so the check measures the production host and a future host/redirect change can't silently mask a regression.

    2. The legacy -> workflow flip is now a production-domain risk, not just a github.io one. The domain root currently is the committed docs/ folder (that's why openamer.com/ -> 200 with the hand-written landing while openamer.com/docs/ -> 404). The "Stage deployment" step (per the 09:07Z analysis) only produces _site/docs/* + llms.txt - there is no _site/index.html. So flipping Pages to "GitHub Actions" before the 09:07Z staging fix would 404 the marketing root on openamer.com itself, not merely on github.io. The staged root pages (index.html, superiority.html, academy.html, consulting.html, skills-store.html, wiki/) must land in _site/ first.

    Net: action items are unchanged from 07:04Z/09:07Z - merge #148 (still open, build_type still legacy), flip the Pages source, stage the root pages - plus one new one: retarget the freshness probe to openamer.com. The legacy source mode is what keeps /docs/** 404 on the new domain; the domain change alone does not touch that.

  16. openamer commented on Oct 11, 2026

    @openamer
    Owner

    Live re-measure at 2026-10-11 ~19:20 UTC + the actual content of the "6/8 Python slices red" blocker on #148 — it is repo-wide red, not #148's diff.

    Unauthenticated curl -L right now:

    URL Result
    https://openamer.github.io/openamer/ 200
    https://openamer.github.io/openamer/docs/ 404
    https://openamer.github.io/openamer/docs/api/skills-index.json 404
    https://openamer.com/docs/ 404
    https://openamer.com/docs/api/skills-index.json 404

    GET /repos/openamer/openamer/pages -> build_type: legacy, source: main:/docs, status: building. A pages build and deployment run completed success at 19:07:54Z (run 38166709836) and /docs/** is still 404. Same root cause as the 07:04Z / 09:07Z checks: in legacy mode the build publishes the contents of docs/ at the domain root (openamer.com/ = the committed landing, 200), so a /docs/ path cannot exist. Nothing on the merge path has changed.

    New this run — I opened the failing slice logs instead of re-asserting "6/8 red". On #148's head (d4a7d08) and on a current-main run (38165823728, 18:55Z) the slices fail on pre-existing, unrelated tests — none of them touch what #148 changes:

    • 2/8 — test_feishu_load_settings_ignores_extra_allow_bots (assert 'all' == 'none'), test_default_compact_banner_keeps_legacy_openamer_openamer_branding, test_kanban_db_has_no_security_findings_if_present
    • 4/8 — tests/test_code_review_bot.py (3), tests/tools/test_video_generation_tool_surface_matrix.py xai edit/extend (4)
    • 5/8 — test_auth_xai_oauth_provider, test_swarm_orchestrator (2), tests/tools/test_asi_darwin.py (3)
    • 6/8 — tests/scripts/test_scoreboard.py::test_every_capability_value_names_its_source
    • 7/8 — test_auth_codex_quota_probe.py (3), test_autonom_watchtower, test_test_suite_watchdog.py (6)
    • 8/8 — test_darwin_engine, test_swarm_home_consistency (2), test_image_generation_env (3), test_subprocess_stdin_guard, test_windows_native_support

    That is ~35 distinct failures across ~14 files. The current-main run fails the same slices (2,3,4,5,7,8), so main is red and #148 merely inherits it. Conclusion: "merge #148" needs those ~35 tests green first (or a maintainer decision to relax the required check) — the ClawHub-walk fix in #148 is necessary but not sufficient on its own.

    One tempting-looking shortcut worth ruling out: #157 pins PYTHONUTF8=1 in the hermetic runner and fixes the Windows canonical-suite red (507 failing locally). That is a different red from these Linux CI slices — the failures above reproduce on ubuntu-latest, where utf8_mode is already 1. #157 is a good fix, but by itself it will not clear #148's slice failures (its own CI run 38161062578 is red on the same slices).

    Net action list: (1) clear the ~35 pre-existing slice failures (new item this run); (2) merge #148; (3) flip Pages to build_type: workflow only after the 09:07Z root-page staging lands; (4) retarget the freshness probe to the canonical openamer.com host. Site build lives under website/.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions