File Types and File Sizes

1 Why this matters

RNA-seq analysis creates many files. Some are small text files that describe your samples. Others are large sequencing or alignment files that can use a lot of space.

This page is a gentle introduction to the file types you may see during RNA-seq analysis. You do not need to understand every detail yet. The goal is to become familiar with the names and what they are generally used for.

Before running a pipeline, it helps to recognize which files are raw data, which files describe the reference genome, and which files are results.

2 Common file types

File type Example What it is Where it comes from Once or per dataset?
FASTQ sample_R1.fastq.gz Raw sequencing reads Sequencing facility Each dataset
CSV samplesheet.csv Table describing samples and input files You create it Each dataset
FASTA genome.fa Reference genome or transcript sequences Downloaded Once, reused
GTF/GFF annotation.gtf Gene annotation file Downloaded Once, reused
BAM sample.bam Aligned reads Pipeline output Each dataset
BAI sample.bam.bai Index for a BAM file Pipeline output Each dataset
TSV/CSV salmon.merged.gene_counts.tsv Count or expression table Pipeline output Each dataset
HTML multiqc_report.html Interactive report opened in a browser Pipeline output Each dataset
LOG/TXT .nextflow.log Logs and command output Pipeline output Each dataset

3 What file extensions mean

A file extension is the short suffix after the last dot in a filename. It tells you what format the file is in, or what program created it.

Extension What it is
.sh Shell script: a text file containing terminal commands that run in order
.sbatch Same as .sh: just an alternate name some people use for SLURM job scripts
.csv Comma-separated values: a table where commas separate columns (open in Excel or R)
.tsv Tab-separated values: same idea, but tabs separate columns instead of commas
.json Structured settings file using key-value pairs: used to pass parameters to Nextflow
.txt, .log, .out, .err Plain text files: usually log messages and error output
.fastq or .fq Raw sequencing reads
.fastq.gz Same, but compressed with gzip: smaller on disk; most tools read .gz directly
.fa or .fasta FASTA format: genome sequence or transcript sequences
.gtf or .gff Gene annotation file: describes where genes are in the genome
.bam Binary alignment file: aligned reads (not human-readable; needs special tools)
.bai Index for a BAM file: required to quickly look up specific positions in a BAM
.html Web page: open in a browser (used for MultiQC reports)
.R, .Rmd, .qmd R script, R Markdown, or Quarto document
Note

Some files have two-part extensions. .fastq.gz means the file is in FASTQ format and it has been compressed with gzip. The .gz part does not change what is inside, it just makes the file smaller for storage and transfer.

4 FASTQ files

FASTQ files are usually the main input for RNA-seq pipelines.

Paired-end data often has two files per sample:

sample1_R1.fastq.gz
sample1_R2.fastq.gz

4.1 Anatomy of a FASTQ filename

Illumina sequencers produce FASTQ files with names that follow a consistent pattern. Here is an example from the S. parvus dataset:

C45-1B_S2_L002_R1_001.fastq.gz

Each part of the name means something:

Part Example Meaning
Sample name C45-1B The sample label set during sequencing. In this lab: condition-timepoint-replicate-tissue (B = body, H = head)
Illumina sample number S2 Assigned automatically during demultiplexing: which barcode belonged to which sample
Lane L002 Which physical lane on the sequencing flowcell was used. Some runs use multiple lanes
Read direction R1 Whether this is read 1 or read 2 of a paired-end run
File number 001 Usually 001; Illumina sometimes splits one sample across multiple files
Format .fastq.gz FASTQ format, compressed with gzip

So C45-1B_S2_L002_R1_001.fastq.gz is: sample C45-1B (body, replicate 1), demultiplexed as sample #2, from lane 2, read 1.

Its paired file would be C45-1B_S2_L002_R2_001.fastq.gz.

FASTQ files can be large. A single compressed FASTQ file may be several gigabytes.

5 Reference files

Reference files describe the genome or transcriptome you want to compare your reads against.

Common reference files include:

genome.fa
annotation.gtf

For non-model organisms like Staurois parvus, reference files may come from a lab assembly or a closely related species, depending on the project.

Note

The choice of reference affects the interpretation of RNA-seq results. Always record which genome and annotation files were used.

6 Pipeline outputs

nf-core pipelines create organized output folders. The exact files depend on the options you choose, but common outputs include:

  • Quality-control reports
  • Quantification tables
  • Alignment files
  • Pipeline logs
  • Software version files
  • MultiQC summary reports

The most useful output to open first is usually:

multiqc_report.html

7 Intermediate files

RNA-seq pipelines also create intermediate files while they are running.

Intermediate files are files created between the raw input files and the final results. They are useful for the pipeline, but they are not always files you need to inspect directly.

One common example is the Nextflow work/ directory. This directory can become large because it stores temporary files for many pipeline steps.

Note

For now, just know that intermediate files can take up space. Later pages will show how to check file sizes and decide what can be cleaned up.

8 Typical storage needs

Exact storage needs depend on the number of samples, read depth, reference files, and pipeline options.

As a rough guide:

Project size Input FASTQs Suggested scratch space
Small test run Less than 5 GB 30 GB
Small project 10-50 GB 100-300 GB
Medium project 50-200 GB 300 GB-1 TB
Large project More than 200 GB 1 TB or more

For this tutorial, we create a workspace with 30 GB because it is meant for a small demo.

9 What files are most important?

Usually keep:

  • samplesheet.csv
  • Final results folder
  • MultiQC report
  • Count or expression tables
  • Pipeline information files
  • Exact commands and parameters used

Usually do not keep forever:

  • Large temporary work/ files
  • Failed test outputs
  • Duplicate downloaded references
  • Old intermediate files that can be regenerated

10 Big picture

  • Raw FASTQ files are the original sequencing data.
  • Reference files tell the pipeline what genome or annotation to use.
  • Output files are the results and reports from the pipeline.
  • Intermediate files help the pipeline run but may be large.
  • Metadata and notes make the analysis reproducible.
Tip

When in doubt, do not delete raw data, metadata, or final results. Temporary and intermediate files are easier to regenerate than original data.