Interpret Results
1 Did the pipeline succeed?
The first thing to check is whether the pipeline actually finished without errors.
In your output log file, a successful run ends with:
-[nf-core/rnaseq] Pipeline completed successfully -
If you see Pipeline completed with errors or the log cuts off mid-run, something went wrong. See the Troubleshooting page.
2 What output folders were created?
After a successful run, your results directory will contain:
| Folder | What it contains |
|---|---|
fastqc/ |
Per-sample read quality reports (before and after trimming) |
trimgalore/ |
Trimming statistics and reports |
star_salmon/ |
Alignment files (BAM), transcript quantification (Salmon), and count tables |
multiqc/ |
A single HTML report summarizing quality across all samples |
pipeline_info/ |
Software versions, parameters used, execution report |
The most important files for you are:
multiqc/star_salmon/multiqc_report.html, your first stop for QCstar_salmon/salmon.merged.gene_counts.tsv, the count table you will use for DESeq2
3 Get the files to your laptop
The MultiQC report is an HTML file that you open in a browser. Copy it from Unity to your laptop using scp. Run this command from a terminal on your laptop (not from Unity):
scp kfloer_smith_edu@unity.rc.umass.edu:/scratch4/workspace/kfloer_smith_edu-YOUR_WORKSPACE/results/multiqc/star_salmon/multiqc_report.html ~/Desktop/Then open it by double-clicking it, or from the terminal:
open ~/Desktop/multiqc_report.htmlTo copy the entire results folder (for archiving):
scp -r kfloer_smith_edu@unity.rc.umass.edu:/scratch4/workspace/kfloer_smith_edu-YOUR_WORKSPACE/results/ ~/Desktop/rnaseq_results/4 Reading the MultiQC report
MultiQC combines results from many tools into one scrollable HTML page. Here is what to look at and what to look for.
4.1 General statistics table
This is at the top of the report. It gives a quick overview of every sample. Columns to check:
| Column | What it means | What to look for |
|---|---|---|
% Dups |
Percentage of duplicate reads | Under 50% is typical; very high duplication can indicate over-amplification |
% GC |
GC content of reads | Should be similar across samples; big differences can indicate contamination |
M Reads |
Millions of reads per sample | Check that all samples have comparable read counts; a very low count is a flag |
4.2 FastQC: sequence quality
Look at the Per Base Sequence Quality plot. This shows quality scores along the length of each read.
- Quality above 30 (the green zone) across most of the read is good
- A drop at the ends is normal, Illumina reads often drop in quality toward the 3’ end
- If quality drops early (before position 50), check whether trimming helped
4.3 FastQC: adapter content
Adapter contamination means sequencing read through the end of a fragment and into the Illumina adapter sequence. nf-core/rnaseq trims adapters automatically with Trim Galore, so you should see minimal adapters in the post-trimming FastQC plot.
- The post-trimming plot should show adapter content close to 0%
- If it is still high after trimming, the adapter detection may have failed, check the library prep documentation
4.4 STAR: alignment rate
The STAR section shows what percentage of reads mapped to the genome.
| Alignment rate | Interpretation |
|---|---|
| 85–95% | Good |
| 70–85% | Acceptable: check for rRNA contamination or genome completeness |
| Below 70% | Investigate: possible wrong reference, contamination, or library prep issue |
For S. parvus, alignment rates may be lower than for model organisms because the genome assembly is less complete. A rate in the 70–85% range is not unusual.
Also look at:
- Uniquely mapped, reads that map to exactly one place in the genome (best for quantification)
- Multi-mapped, reads that map to more than one place (less useful; often from repetitive regions)
4.5 Salmon: quantification
Salmon uses the aligned reads to estimate transcript abundances. The Salmon section shows the percentage of reads assigned to genes (mapping rate).
- A Salmon mapping rate similar to the STAR alignment rate is expected
- If Salmon’s rate is much lower than STAR’s, check whether the GTF matches the genome FASTA
4.6 Gene body coverage (RSeQC)
This plot shows read coverage across the length of genes, from 5’ (left) to 3’ (right).
- A flat or slightly 5’-biased curve is ideal
- A strong 3’ bias suggests RNA degradation in the sample, poor RNA quality at time of extraction
- A strong 5’ bias is rarer and may indicate specific library prep issues
4.7 Sample clustering and correlation
MultiQC often shows a sample-to-sample correlation heatmap or PCA from DESeq2. Use this to check:
- Do biological replicates cluster together?
- Do samples from the same tissue or treatment group cluster together?
- Is there an obvious outlier sample?
If a sample looks very different from everything else, check its alignment rate, read count, and gene body coverage before including it in downstream analysis.
5 Understanding the count table
The key output for downstream analysis is:
star_salmon/salmon.merged.gene_counts.tsv
This is a tab-separated table where:
- Each row is a gene
- Each column is a sample
- Each cell contains the estimated count of reads for that gene in that sample
These counts are what you pass to DESeq2 (or another differential expression tool) to find genes that differ between conditions.
The counts are raw estimated counts, not normalized values. DESeq2 handles normalization internally. Do not normalize the counts yourself before running DESeq2.
You can preview the file on Unity:
head star_salmon/salmon.merged.gene_counts.tsvCopy it to your laptop for use in R:
# Run on your laptop
scp kfloer_smith_edu@unity.rc.umass.edu:/scratch4/workspace/kfloer_smith_edu-YOUR_WORKSPACE/results/star_salmon/salmon.merged.gene_counts.tsv ~/Desktop/6 Save provenance files
Before you clean up the workspace or move on, archive these files, they document exactly what was run and make the analysis reproducible:
pipeline_info/software_versions.yml, every tool version usedpipeline_info/execution_report.html, a visual timeline of all pipeline stepspipeline_info/execution_timeline.html, timing informationpipeline_info/params_*.json, the exact parameters used
# Copy the pipeline_info folder to lab storage
cp -r results/pipeline_info/ /work/pi_lmangiamele_smith_edu/YOUR_PROJECT_NAME/pipeline_info/7 What to do next
Once you are satisfied with the quality of the run:
- Copy the count table (
salmon.merged.gene_counts.tsv) and thetx2genefile to a safe location - Copy the MultiQC report and pipeline_info folder for your records
- Proceed to Differential Gene Expression Analysis to run the DESeq2 analysis in R