Build a local database of (almost) every public, non-fork GitHub repository with at least 10 stars, as the data foundation for a visualization project.
The exact target — current, all non-fork repos with stars ≥ 10 — can't be served by a single source, so this project combines two:
| Source | Role | Why |
|---|---|---|
| GitHub Search API | Enumerate the repo list | The only source that natively filters stars:>=10 fork:false on current data. |
| ecosyste.ms | Enrich each repo | Rich per-repo metadata (language, dependencies, SBOM, topics, license). |
- GitHub coverage on ecosyste.ms: ~286M repos, kept in sync (its live API is current).
- But the ecosyste.ms list API can't enumerate by stars: no star filter, any
sorton the host-level/repositoriesendpoint returns 500, and pagination is hard-capped at 100 pages (max 100k rows/query). - Its downloadable data dump (
/open-data) is a one-off snapshot frozen at 2023-08-30 (211 GB, ~168M repos) — too stale for a current view.
So ecosyste.ms is used for enrichment, not enumeration.
GitHub Search returns at most 1000 results per query, but stars:>=10 fork:false
matches ~2.75M repos (and a single bucket like stars:10 alone is ~200k). The
crawler walks the keyspace by recursively bisecting stars × created-date until
every leaf window holds ≤ 1000 results, guaranteeing full coverage without relying on
sort.
- Language: Go
- Storage: DuckDB (single-file, columnar — fast group-by/ranking for visualization)
A GitHub token is required for the Search API. The crawler resolves it from, in order:
GITHUB_TOKEN → GH_TOKEN → gh auth token.
enrich joins the ecosyste.ms polite pool by sending a contact email as the
mailto= query parameter, which raises the rate limit from ~5000 to ~15000 req/hr
(≈3× faster). Set the email with -mailto or ECOSYSTEMS_MAILTO; pass -mailto ""
to stay anonymous. Note: only the mailto= query parameter engages the polite pool
— putting the email in the User-Agent does not (verified 2026-06).
# Build (CGO required for the DuckDB driver)
CGO_ENABLED=1 go build -o bin/crawler ./cmd/crawler
# Verify the GitHub token resolves, then create the database
./bin/crawler token-check
./bin/crawler init-db
# 1) Enumerate via GitHub Search (resumable — re-run to continue after Ctrl-C).
# Full run over ~2.75M repos is rate-limited to 30 req/min ⇒ many hours;
# Ctrl-C any time and re-run to resume from where it stopped.
./bin/crawler enumerate # stars >= 10, all dates (the full target)
./bin/crawler enumerate -min-stars 1000 # a smaller slice to start with
# 2) Enrich each stored repo with ecosyste.ms metadata (resumable).
./bin/crawler enrich # all pending
./bin/crawler enrich -limit 500 # just the first 500 pending (highest-star first)
# 3) Refresh the breakout board's rolling 7-day window (seconds, a few requests).
./bin/crawler breakout # new repos with >= 300 stars, last 7 days
./bin/crawler breakout -window 336h # a wider window
# Inspect progress at any time
./bin/crawler statsDB_PATH overrides the DuckDB file (default data/repos.duckdb).
- enumerate: every fully-drained
(stars × created-date)leaf window is recorded incrawl_windows; a re-run skips windows already done. Window bounds are deterministic (fixed star ceiling), so a re-run is an exact no-op once complete. - enrich: repos are processed in
eco_synced_at IS NULLorder; each processed repo (including 404s) is stamped, so a re-run only picks up the remainder. - breakout keeps no checkpoints at all — it is a rolling snapshot, so every run re-fetches the whole window and refreshes the star counts in place.
A repository created N days ago necessarily started at zero stars. So "went
from nothing to 1000+ stars in a week" needs no star history: created_at
plus the current star count fully define the set, and one Search query
(stars:>=300 fork:false created:<7d ago>..<today>) returns all of it.
Measured June–August 2026, a week holds 20–35 repos past 1000 stars and ~90–115 past 300 — scarce, but never empty.
This deliberately does not ride on enumerate. Enumerate walks the whole
keyspace once and checkpoints each drained window, so a repo created today is
not guaranteed to surface until that window is walked again — fine for a
census, useless for a board that must be current. breakout is a handful of
requests over a fixed 7-day window instead, cheap enough to re-run hourly.
Repos below 1000 stars but above the 300 fetch floor become the "still climbing" tier on the site. The window rolls: once a repo is older than seven days it leaves the board, which is why nothing needs to be archived.
repos — one row per repository (GitHub fields + eco_* enrichment columns).
crawl_windows — enumeration checkpoints for resumability.
Query the DuckDB file directly for visualization, e.g.:
SELECT language, count(*) AS repos, sum(stars) AS stars
FROM repos GROUP BY language ORDER BY stars DESC LIMIT 20;.github/workflows/crawl.yml runs the whole crawl in CI so you don't have to keep
a local machine online. Each run:
- restores the DuckDB from the
dbGitHub Release (assetrepos.duckdb.gz), - builds the crawler,
- runs
breakout(seconds), thenenumerate+enrichbounded by-max-runtime, - gzips and re-uploads the DB to the release (
gh release upload --clobber).
breakout runs before the long crawl so a run that exhausts its time budget
still leaves a current board. It is not a separate daily workflow on purpose:
the DuckDB is single shared state behind one release asset, so a second
scheduled workflow would race this one and clobber the upload.
An hourly cron chains runs; resumability means each picks up where the last
stopped. Rate-limit quotas reset hourly, so we trigger hourly — but the
concurrency guard keeps runs from overlapping: an hourly trigger that fires
while a long job is still running just queues and starts the moment that job
ends. The net effect is near-continuous ~5.5h runs (the DB is transferred once
per run, not once per hour) that still consume each hour's quota. The db
release is created automatically on the first run. Trigger manually from the
Actions tab ("Run workflow") to tune the inputs:
| Input | Default | Meaning |
|---|---|---|
min_stars |
10 |
lower star bound for enumerate |
enumerate_minutes |
180 |
time budget for enumerate (0 to skip) |
enrich_minutes |
120 |
time budget for enrich (0 to skip) |
mailto |
contact email | ecosyste.ms polite-pool address |
Notes:
- Use a public repo. Actions standard runners are free with no monthly cap on public repos; private repos are limited to 2000 min/month — far too little for the multi-day full crawl.
- The full DB approaches the 2 GB per-asset release limit;
repos.duckdb.gzkeeps headroom. If it ever exceeds 2 GB, switch the workflow to external object storage. - Once enumeration is complete, set
enumerate_minutes=0on the scheduled run to avoid re-counting the internal window tree each time, leaving the full budget for enrich. - Per-job time is capped at 6h;
-max-runtimemakes each phase exit cleanly before then with progress saved.
web/ is a static visualization site (Vite + TypeScript, hand-built SVG — no chart
library). It reads small pre-aggregated JSON for the charts, so the browser never
loads the full database for those. The landing section is the breakout board
(repos that went 0 → 1000+ stars this week, plus a still-climbing tier); below it
sit star ranking, language distribution, growth over time (by creation
year), a crawl coverage map, project health, and a topic constellation.
The ranking search covers the entire dataset, not just the top repos: the
aggregate step also exports repos.parquet (every repo, columns the ranking needs,
stars-DESC, ZSTD), and the browser queries it with DuckDB-WASM over HTTP range
requests — only the byte ranges a query touches are fetched, so search spans all
~millions of repos with no backend. First paint is instant from top_repos.json;
if DuckDB-WASM is unavailable the search degrades to filtering that top-1000 set.
The data pipeline is crawler aggregate, which turns the DuckDB into a handful of
small JSON files plus repos.parquet:
./bin/crawler aggregate -out web/public/data # meta/top_repos/trends/topics/breakout.json + repos.parquet
cd web && pnpm install && pnpm dev # local dev.github/workflows/pages.yml rebuilds and deploys the site to GitHub Pages
after each crawl finishes (and on demand): it restores the DB from the db
release, runs aggregate, builds web/, and deploys. Enable Pages once with
source = GitHub Actions (Settings → Pages). The site lives at
https://hoveychen.github.io/opensource-world/.