A self-maintaining, token-minimal library of nf-core pipelines for AI agents.
Each pipeline is a git submodule plus an auto-generated skill.md an agent reads to run it
correctly — no external lookups, no hallucinated flags.
🌐 Live site: danilomonge.github.io/nf-claw — a data-driven interface to the whole library (pipelines, parameters, docs and automation), rebuilt from the repository on every push.
pipelines/<name>/—upstream/(submodule, pinned to a release) + generatedskill.md(run command, inputs and required parameters with their allowed values and constraints, plus a map of parameter groups) andreference.md(every parameter, with required/hidden flags, allowed values, value constraints and default)runner/— thenfclawruntime (run a pipeline); the agent's toollibrarian/— builds the context files and bumps submodules (run viamake)catalog.md/catalog.json— the index of available pipelines, with each one'sinput(derived from its samplesheet schema),output(the guaranteed output contract),summary(the authors' own one-paragraph description from the pipeline README; inskill.mdandcatalog.json) andtools(the methods it runs, from the pipeline's ownCITATIONS.md)sources.tsv— the source list (name, url, version policy)
Design details (the three zones and how generation stays drift-free): docs/architecture.md.
Version/engine compatibility (DSL2-only; each release runs with the Nextflow version it declares): docs/compatibility.md.
Run-time errors (space-in-path, IPv6 host, no-network, Nextflow-version parse errors, known upstream-pipeline bugs): docs/known-issues.md.
pip install -e . # first time: puts `nfclaw` on PATH (or use `python3 -m runner <cmd>`)
nfclaw list # or: python3 -m runner list
nfclaw run rnaseq --input samplesheet.csv --outdir results -profile docker
# run a specific (non-latest) release — default stays the pinned latest
nfclaw versions rnaseq # list release tags (latest is flagged)
nfclaw show rnaseq --pipeline-version 3.14.0 # that release's skill.md (schema + flags), generated on demand
nfclaw run rnaseq --pipeline-version 3.14.0 --input samplesheet.csv --outdir results -profile docker
# pin the Nextflow engine / set NXF_* for one run (both recorded in provenance)
nfclaw run rnaseq --nxf-ver 25.10.2 --input ss.csv --outdir results -profile docker # if a newer Nextflow breaks the release
nfclaw run rnaseq --nxf-env NXF_JVM_ARGS=-Djava.net.preferIPv6Addresses=true ... # IPv6-only host (JVM → GitHub)
# cap what the run may request, for a machine smaller than nf-core's server-sized defaults
nfclaw run rnaseq --input ss.csv --outdir results -profile docker \
--limit-cpus 4 --limit-memory 15.GB --limit-time 1.h # Nextflow process.resourceLimitsThe editable install is intentional: nfclaw reads pinned submodules and generated context from
this checkout. A standalone wheel contains the Python runtime but not the pipeline library, so the
CLI exits with a setup error instead of silently reporting an empty catalog.
make build # regenerate skill.md/reference.md/catalog from the submodules
make update # bump submodules to the latest release tag, then rebuild
make check # drift gate + tests
make testAppend a line to sources.tsv (name<TAB>url<TAB>latest-release), then:
git submodule add <url> pipelines/<name>/upstream
git -C pipelines/<name>/upstream checkout tags/<X.Y.Z> # pin to a release (not the default branch)
make buildFull step-by-step, with the version-policy column and the checks to run, is in CONTRIBUTING.md.
Five workflows keep the library — and the site — up to date with no manual edits:
auto-update.yml(daily): finds each pipeline's newest release withgit ls-remote --tags(pure git, no APIs), checks it out, regenerates context, and opens a PR. Updated pipelines must pass unit tests, the drift gate, and Nextflow acceptance in the same job before the PR is merged automatically, which triggers a site rebuild.discover-pipelines.yml(weekly): finds DSL2 nf-core pipelines not yet insources.tsv, scaffolds each one (pinned submodule + generated context), then validates the batch — unit tests, the drift gate and Nextflow acceptance (nextflow -preview). Any pipeline Nextflow rejects is dropped; if the rest pass, the PR is opened and auto-merged (which rebuilds the site).deploy-pages.yml(on every push tomain): builds the website as a static export and publishes it to GitHub Pages, so the live site always reflects the current repository state.smoke.yml(weekly + on pipeline/runner changes): for every pipeline, builds and preflights the demo command (nfclaw run --check --demo) — parsing the pinned schema, validating parameters and composing the test profile, all without executing the workflow. It guards what nf-claw owns (schema parse, generated command, submodule integrity).nextflow-validate.yml(nightly + on pipeline changes): runs each pipeline through Nextflow with-profile test,docker -preview, so Nextflow compiles it, resolves the config/profile, runs the nf-schema parameter validation and builds the full task DAG — then stops before launching any task. This proves every pinned pipeline is accepted and runnable up to the point of execution; the analysis itself is already covered by nf-core's own CI before a release is tagged.
The drift gate guarantees committed context always matches the pinned submodule. More detail:
docs/updating.md.
The pipelines themselves are the work of the nf-core community — the heart of the library — and are wrapped unmodified as pinned git submodules; each keeps its own authors, license and citation. nf-claw (the wrapper/runtime) was created by Danilo Monge (Eberhard Karls Universität Tübingen), and adapts some files and its repository structure from ClawBio — created by Manuel Corpas (MIT; copyright retained in LICENSE, provenance in NOTICE).
If you use nf-claw, cite it via CITATION.cff, together with the tools it builds on:
- Nextflow (workflow engine) — Di Tommaso P, et al. Nextflow enables reproducible computational workflows. Nat Biotechnol 35, 316–319 (2017). doi:10.1038/nbt.3820
- nf-core (the original pipelines) — Ewels PA, et al. The nf-core framework for community-curated bioinformatics pipelines. Nat Biotechnol 38, 276–278 (2020). doi:10.1038/s41587-020-0439-x
- the specific pipeline you ran — each lists its own reference in
pipelines/<name>/upstream/CITATIONS.md - ClawBio (the predecessor nf-claw adapts files and structure from, by Manuel Corpas) — Corpas M. ClawBio: Bioinformatics-Native AI Agent Skill Library. Zenodo (2026). doi:10.5281/zenodo.19420648
nf-claw is MIT-licensed and includes parts adapted from ClawBio (© Manuel Corpas, MIT — see LICENSE and NOTICE); the nf-core pipelines it wraps are MIT, and Nextflow is Apache-2.0.
Requires: Python 3.11+, git, Nextflow (Java 17+), Docker/Singularity.