Data Archiving and Sharing
Raw sequencing data, genome assemblies, and manuscript-linked code should be deposited in the appropriate specialist or general repository.
- GSA Raw metagenomic, amplicon, transcriptomic, and other sequence data with BioProject and BioSample metadata.
- SRA / ENA International raw-sequence archiving and public data discovery; the same dataset does not need duplicate submission to multiple INSDC nodes.
- GWH Genome assemblies, annotations, and associated metadata for isolates and MAGs.
- Zenodo Versioned archives for figure data, scripts, supplementary files, and software, with DOI assignment.
Microbial Annotation and Functional Search
Taxonomic and functional interpretation requires the reference release, download date, and key settings to be recorded.
- SILVA rRNA alignment and taxonomy for bacteria, archaea, and eukaryotes; useful for 16S analysis and classifier training.
- UNITE Fungal ITS taxonomy and Species Hypothesis reference sequences.
- GTDB Genome-phylogeny-based bacterial and archaeal taxonomy for isolates and MAGs.
- eggNOG / KEGG Cross-annotation of protein orthologs, functional categories, metabolic pathways, and modules.
Environmental and Ecological Open Data
Resources for study-area context, ecological observations, and publication-linked environmental datasets.
- National Earth System Science Data Center Chinese climate, soil, land-use, geospatial, and environmental baseline data.
- PANGAEA Long-term publication and citation of geochemical, sediment, marine, and environmental datasets.
- GBIF Global species occurrence, specimen, and observation records; retain the download DOI and filter criteria.
Analysis Environments and Reproducible Workflows
Analyses should be rerunnable rather than documented only by software names or one-off operations.
- QIIME 2 Amplicon analysis, visualization, and provenance; retain primer, denoising, filtering, and classifier settings.
- fastp / MultiQC Raw-read quality control and summary reports; useful as a standard first step in sequencing analysis.
- R / Python Statistics, visualization, automated data processing, and modeling; retain scripts and package versions for reported results.
- Conda, containers, and workflows Environment files, images, and Snakemake / Nextflow preserve dependencies and steps in complex analyses.
Minimum Data Practices
- Assign unique project and sample identifiers at project start, and complete metadata sheets before sampling or sequencing submission.
- Place raw data in read-only locations; process copies rather than overwriting instrument files, FASTQ files, raw images, or spectra.
- Record data sources, software and database versions, parameters, run dates, and reasons for excluding samples.
- Retain source data, scripts, and documentation for each reported figure; do not report statistical results with P values alone.
- Back up important data according to the 3-2-1 principle, and complete repository submission, accession logging, and release-date settings before manuscript submission.