Frog Data Availability

1 Purpose

This page collects the data files and repositories used for Staurois parvus RNA-seq analysis in the Mangiamele Lab.

2 What is SRA, and how does it relate to NCBI?

NCBI (National Center for Biotechnology Information) is a US government database that stores biological data, genome sequences, gene records, publications, and more. It is the main public archive for genomics data.

SRA (Sequence Read Archive) is the part of NCBI that specifically stores raw sequencing reads, the FASTQ files that come off a sequencer. When a paper publishes RNA-seq data, the raw reads are usually deposited in SRA, and each dataset is given a unique accession number that starts with SRA, SRP, SRR, or PRJNA.

To download data from SRA you use a tool called sra-tools (specifically the fasterq-dump command), which converts SRA format back into FASTQ files.

GEO (Gene Expression Omnibus) is a related NCBI database that stores processed expression data, count tables, normalized values, and metadata. GEO records often link back to the raw SRA reads.

In practice: if you see a paper say “data available at GEO accession GSE######”, the raw reads for that study are usually also in SRA under a linked accession.

3 What you’ll most likely need

These are the files you need to run the nf-core/rnaseq pipeline and the downstream DESeq2 analysis.

3.1 RNA-seq reads (flutamide experiment)

Raw paired-end FASTQ files from the flutamide treatment experiment (Kate’s thesis).

On Unity:

/work/pi_lmangiamele_smith_edu/03_26_flut_yale_rnaseq

Yale server (external):

http://fcb.ycga.yale.edu:3010/fTvjY60iTRcfef5EWgR3HlRd5pz98l0/sample_dir_000020575/

3.2 Genome assembly (EGAPx)

The S. parvus genome assembly annotated with EGAPx. Contains FASTA, GTF, and GFF files.

On Unity:

/work/pi_lmangiamele_smith_edu/output_tadpole_plus_adult

Key files in that directory:

File What it is
complete.genomic.fna Genome FASTA (reference sequence)
complete.genomic.gtf Gene annotation in GTF format
complete.genomic.gff Gene annotation in GFF format

4 Some other useful data

These files are not needed for the main RNA-seq pipeline but are useful context for the project.

4.1 HiFi PacBio reads (used to build the EGAPx assembly)

On Unity:

/work/pi_lmangiamele_smith_edu/yale_hifi_03_26

Yale server (external):

http://fcb.ycga.yale.edu:3010/Bcs7fnShP4sYdsw2ZmkH4Fxpcn83c60/20260313_lmangiamele_r84189_20260310_214511/

4.2 Old CLR PacBio reads

On Unity:

/work/pi_lmangiamele_smith_edu/pacbio_parvus

4.3 Public genome resources

Resource Link Notes Year
Old genome assembly GCA_951230385.1 Paper 2023
New genome assembly Add link Add accession once deposited 2026

5 Download notes

Whenever you use a dataset, record:

  • Where you got it (Unity path or external URL)
  • Date accessed
  • File names downloaded
  • Genome or annotation version

This information makes the analysis reproducible and saves time if you need to re-run anything later.