File Types and File Sizes
1 Why this matters
RNA-seq analysis creates many files. Some are small text files that describe your samples. Others are large sequencing or alignment files that can use a lot of space.
This page is a gentle introduction to the file types you may see during RNA-seq analysis. You do not need to understand every detail yet. The goal is to become familiar with the names and what they are generally used for.
Before running a pipeline, it helps to recognize which files are raw data, which files describe the reference genome, and which files are results.
2 Common file types
| File type | Example | What it is | Where it comes from | Once or per dataset? |
|---|---|---|---|---|
| FASTQ | sample_R1.fastq.gz |
Raw sequencing reads | Sequencing facility | Each dataset |
| CSV | samplesheet.csv |
Table describing samples and input files | You create it | Each dataset |
| FASTA | genome.fa |
Reference genome or transcript sequences | Downloaded | Once, reused |
| GTF/GFF | annotation.gtf |
Gene annotation file | Downloaded | Once, reused |
| BAM | sample.bam |
Aligned reads | Pipeline output | Each dataset |
| BAI | sample.bam.bai |
Index for a BAM file | Pipeline output | Each dataset |
| TSV/CSV | salmon.merged.gene_counts.tsv |
Count or expression table | Pipeline output | Each dataset |
| HTML | multiqc_report.html |
Interactive report opened in a browser | Pipeline output | Each dataset |
| LOG/TXT | .nextflow.log |
Logs and command output | Pipeline output | Each dataset |
3 What file extensions mean
A file extension is the short suffix after the last dot in a filename. It tells you what format the file is in, or what program created it.
| Extension | What it is |
|---|---|
.sh |
Shell script: a text file containing terminal commands that run in order |
.sbatch |
Same as .sh: just an alternate name some people use for SLURM job scripts |
.csv |
Comma-separated values: a table where commas separate columns (open in Excel or R) |
.tsv |
Tab-separated values: same idea, but tabs separate columns instead of commas |
.json |
Structured settings file using key-value pairs: used to pass parameters to Nextflow |
.txt, .log, .out, .err |
Plain text files: usually log messages and error output |
.fastq or .fq |
Raw sequencing reads |
.fastq.gz |
Same, but compressed with gzip: smaller on disk; most tools read .gz directly |
.fa or .fasta |
FASTA format: genome sequence or transcript sequences |
.gtf or .gff |
Gene annotation file: describes where genes are in the genome |
.bam |
Binary alignment file: aligned reads (not human-readable; needs special tools) |
.bai |
Index for a BAM file: required to quickly look up specific positions in a BAM |
.html |
Web page: open in a browser (used for MultiQC reports) |
.R, .Rmd, .qmd |
R script, R Markdown, or Quarto document |
Some files have two-part extensions. .fastq.gz means the file is in FASTQ format and it has been compressed with gzip. The .gz part does not change what is inside, it just makes the file smaller for storage and transfer.
4 FASTQ files
FASTQ files are usually the main input for RNA-seq pipelines.
Paired-end data often has two files per sample:
sample1_R1.fastq.gz
sample1_R2.fastq.gz
4.1 Anatomy of a FASTQ filename
Illumina sequencers produce FASTQ files with names that follow a consistent pattern. Here is an example from the S. parvus dataset:
C45-1B_S2_L002_R1_001.fastq.gz
Each part of the name means something:
| Part | Example | Meaning |
|---|---|---|
| Sample name | C45-1B |
The sample label set during sequencing. In this lab: condition-timepoint-replicate-tissue (B = body, H = head) |
| Illumina sample number | S2 |
Assigned automatically during demultiplexing: which barcode belonged to which sample |
| Lane | L002 |
Which physical lane on the sequencing flowcell was used. Some runs use multiple lanes |
| Read direction | R1 |
Whether this is read 1 or read 2 of a paired-end run |
| File number | 001 |
Usually 001; Illumina sometimes splits one sample across multiple files |
| Format | .fastq.gz |
FASTQ format, compressed with gzip |
So C45-1B_S2_L002_R1_001.fastq.gz is: sample C45-1B (body, replicate 1), demultiplexed as sample #2, from lane 2, read 1.
Its paired file would be C45-1B_S2_L002_R2_001.fastq.gz.
FASTQ files can be large. A single compressed FASTQ file may be several gigabytes.
5 Reference files
Reference files describe the genome or transcriptome you want to compare your reads against.
Common reference files include:
genome.fa
annotation.gtf
For non-model organisms like Staurois parvus, reference files may come from a lab assembly or a closely related species, depending on the project.
The choice of reference affects the interpretation of RNA-seq results. Always record which genome and annotation files were used.
6 Pipeline outputs
nf-core pipelines create organized output folders. The exact files depend on the options you choose, but common outputs include:
- Quality-control reports
- Quantification tables
- Alignment files
- Pipeline logs
- Software version files
- MultiQC summary reports
The most useful output to open first is usually:
multiqc_report.html
7 Intermediate files
RNA-seq pipelines also create intermediate files while they are running.
Intermediate files are files created between the raw input files and the final results. They are useful for the pipeline, but they are not always files you need to inspect directly.
One common example is the Nextflow work/ directory. This directory can become large because it stores temporary files for many pipeline steps.
For now, just know that intermediate files can take up space. Later pages will show how to check file sizes and decide what can be cleaned up.
8 Typical storage needs
Exact storage needs depend on the number of samples, read depth, reference files, and pipeline options.
As a rough guide:
| Project size | Input FASTQs | Suggested scratch space |
|---|---|---|
| Small test run | Less than 5 GB | 30 GB |
| Small project | 10-50 GB | 100-300 GB |
| Medium project | 50-200 GB | 300 GB-1 TB |
| Large project | More than 200 GB | 1 TB or more |
For this tutorial, we create a workspace with 30 GB because it is meant for a small demo.
9 What files are most important?
Usually keep:
samplesheet.csv- Final results folder
- MultiQC report
- Count or expression tables
- Pipeline information files
- Exact commands and parameters used
Usually do not keep forever:
- Large temporary
work/files - Failed test outputs
- Duplicate downloaded references
- Old intermediate files that can be regenerated
10 Big picture
- Raw FASTQ files are the original sequencing data.
- Reference files tell the pipeline what genome or annotation to use.
- Output files are the results and reports from the pipeline.
- Intermediate files help the pipeline run but may be large.
- Metadata and notes make the analysis reproducible.
When in doubt, do not delete raw data, metadata, or final results. Temporary and intermediate files are easier to regenerate than original data.