Pre-processing pipeline for Species-scale RNAseq from the NCBI Sequence Read Archive (SRA)
RNAquarium is a Nextflow (DSL2) pipeline for reprocessing public RNAseq datasets at species scale. Starting from raw SRA runs, it filters host reads and produces gene counts and non-host reads (Part I), then assembles and taxonomically classifies the remaining non-host content (Part II).
RNAquarium enables:
- Retrieving and quality-filtering RNAseq runs directly from the SRA
- Filtering host reads by repeated alignment to a host genome — for zebrafish, retaining ~0.7% of input reads (~11 billion of 1.64 trillion across 77,188 runs) as "possible nonhost"
- Producing per-dataset gene counts tables
- Assembling and taxonomically classifying non-host reads, including viruses (Part II: metatranscriptomics)
Explore, visualize, and interact with RNAquarium project data — including via a chatbot — at the RNAquarium Portal.
Please refer to the documentation for full installation, parameters, and pipeline reference.
Part I — Transcriptomic + Filtering
- Inputs and Databases
- Non-host pipeline overview
- Getting Started (installation & usage)
- Parameters
- Pipeline Explanation
- Technical Notes
Part II — Metatranscriptomics
- Metatranscriptome pipeline overview (usage)
- Metatranscriptome Technical Notes
- Metatranscriptome virus-steps Technical Notes
For test-running RNAquarium on a SLURM cluster with a tiny, fully reproducible two-sample zebrafish example refer to the walkthrough.
RNAquarium is developed and maintained by the Computational Biology Platform & Balla Group at the Biohub.
- Research Lead Team: Keir Balla, Duo Peng, Yasin Şenbabaoğlu
- Bioinformatics Team: Yttria Aniseia, Eric Waltari, Max Frank, Gibraan Rahman, Andy Zhou, Yang-Joon Kim, Hejin Huang
- Web Development Team: Leandro Lima, Wellington Rutes
This project follows the Code of Conduct. By participating, you are expected to uphold it.
RNAquarium is released under the BSD 3-Clause License. Copyright (c) Biohub.