Frog Data Availability
1 Purpose
This page collects the data files and repositories used for Staurois parvus RNA-seq analysis in the Mangiamele Lab.
2 What is SRA, and how does it relate to NCBI?
NCBI (National Center for Biotechnology Information) is a US government database that stores biological data, genome sequences, gene records, publications, and more. It is the main public archive for genomics data.
SRA (Sequence Read Archive) is the part of NCBI that specifically stores raw sequencing reads, the FASTQ files that come off a sequencer. When a paper publishes RNA-seq data, the raw reads are usually deposited in SRA, and each dataset is given a unique accession number that starts with SRA, SRP, SRR, or PRJNA.
To download data from SRA you use a tool called sra-tools (specifically the fasterq-dump command), which converts SRA format back into FASTQ files.
GEO (Gene Expression Omnibus) is a related NCBI database that stores processed expression data, count tables, normalized values, and metadata. GEO records often link back to the raw SRA reads.
In practice: if you see a paper say “data available at GEO accession GSE######”, the raw reads for that study are usually also in SRA under a linked accession.
3 What you’ll most likely need
These are the files you need to run the nf-core/rnaseq pipeline and the downstream DESeq2 analysis.
3.1 RNA-seq reads (flutamide experiment)
Raw paired-end FASTQ files from the flutamide treatment experiment (Kate’s thesis).
On Unity:
/work/pi_lmangiamele_smith_edu/03_26_flut_yale_rnaseq
Yale server (external):
http://fcb.ycga.yale.edu:3010/fTvjY60iTRcfef5EWgR3HlRd5pz98l0/sample_dir_000020575/
3.2 Genome assembly (EGAPx)
The S. parvus genome assembly annotated with EGAPx. Contains FASTA, GTF, and GFF files.
On Unity:
/work/pi_lmangiamele_smith_edu/output_tadpole_plus_adult
Key files in that directory:
| File | What it is |
|---|---|
complete.genomic.fna |
Genome FASTA (reference sequence) |
complete.genomic.gtf |
Gene annotation in GTF format |
complete.genomic.gff |
Gene annotation in GFF format |
4 Some other useful data
These files are not needed for the main RNA-seq pipeline but are useful context for the project.
4.1 HiFi PacBio reads (used to build the EGAPx assembly)
On Unity:
/work/pi_lmangiamele_smith_edu/yale_hifi_03_26
Yale server (external):
4.2 Old CLR PacBio reads
On Unity:
/work/pi_lmangiamele_smith_edu/pacbio_parvus
4.3 Public genome resources
| Resource | Link | Notes | Year |
|---|---|---|---|
| Old genome assembly | GCA_951230385.1 | Paper | 2023 |
| New genome assembly | Add link | Add accession once deposited | 2026 |
5 Download notes
Whenever you use a dataset, record:
- Where you got it (Unity path or external URL)
- Date accessed
- File names downloaded
- Genome or annotation version
This information makes the analysis reproducible and saves time if you need to re-run anything later.