Troubleshooting
1 Start here: how to find out what went wrong
1.1 Check whether the job is still running
squeue --meIf the job is no longer listed, it has finished, either successfully or with an error.
1.2 Read the output log
tail -n 50 ~/rnaseq_nf_core/job-logs/nfcore-rnaseq-test.JOBID.outReplace JOBID with your actual job ID. The end of the output log usually shows what the pipeline was doing when it stopped.
1.3 Read the error log
cat ~/rnaseq_nf_core/job-logs/nfcore-rnaseq-test.JOBID.errThis file captures error messages from the pipeline. Look for lines starting with ERROR or WARN.
1.4 Find failed task logs inside the work directory
When a pipeline step fails, Nextflow creates a detailed log inside the work/ directory. The output log will print something like:
ERROR ~ Error executing process 'NFCORE_RNASEQ:RNASEQ:...'
...
Work dir:
/scratch4/workspace/kfloer_smith_edu-demo/work/ab/c1234567890abc...
Navigate to that directory and look at the log files:
cd /scratch4/workspace/kfloer_smith_edu-demo/work/ab/c1234567890abc
ls -la
cat .command.err
cat .command.out2 Common problems and fixes
2.1 The job was cancelled: time limit exceeded
Symptom: The job disappears from squeue before finishing. The log ends without a completion message. SLURM emails you “CANCELLED” or “TIMEOUT”.
Fix: Increase --time in the #SBATCH header. For a full dataset, use at least 3-00:00:00 (3 days).
#SBATCH --time=3-00:00:00Then resubmit. The -resume flag in the script will skip steps that already finished.
2.2 The job failed: out of memory
Symptom: The error log contains Out of memory, Killed, or a process exits with code 137.
Fix: Increase --mem in the #SBATCH header:
#SBATCH --mem=128GFor large datasets with many samples, 128G or more may be needed.
2.3 Nextflow module not found
Symptom:
module: command not found
or
ERROR: Unable to find module nextflow/26.04.1
Fix: Check what version is available:
module avail nextflowUpdate the module name in your script to match an available version. You can also browse available modules at the Unity module explorer.
2.4 Apptainer container failed to pull
Symptom: The log shows something like:
Failed to pull singularity image
FATAL: Unable to pull docker://...
Fix: This usually means the container download was interrupted. The script already includes a line to clean up partial downloads before running:
find "$NXF_APPTAINER_CACHEDIR" -type f -name "*.pulling.*" -delete || trueIf the problem persists:
Delete the cache folder and try again:
rm -rf /scratch4/workspace/YOUR_USERNAME-YOUR_WORKSPACE/.nextflow-apptainer-cacheResubmit the job, Nextflow will re-download the containers.
2.5 The samplesheet is not found
Symptom:
ERROR ~ No such file or directory: samplesheet.csv
Fix: Check that the path in your params.json exactly matches where the file actually is:
ls /scratch4/workspace/YOUR_USERNAME-YOUR_WORKSPACE/samplesheet.csvIf the file is missing, create it (see the Run the Pipeline page).
2.6 FASTQ files are not found
Symptom:
ERROR ~ File does not exist: /path/to/sample_R1.fastq.gz
Fix: Check that the paths in your samplesheet are correct and that the files exist:
ls /work/pi_lmangiamele_smith_edu/03_26_flut_yale_rnaseq/Make sure you are using full absolute paths (starting with /) in the samplesheet, not relative paths.
2.7 Workspace ran out of space
Symptom: The log shows No space left on device or a step fails unexpectedly.
Fix: Check how much space is left:
df -h /scratch4/workspace/YOUR_USERNAME-YOUR_WORKSPACE/The work/ folder (where Nextflow stores intermediate files) can grow very large. To remove completed intermediate files while keeping final results, run this from inside the workspace directory:
nextflow clean -f2.8 A pipeline step failed but I do not know why
Steps to diagnose:
- Find the
Work dir:path printed in the error message in your output log - Navigate to that directory:
cd /path/to/work/dir - Read the task’s error log:
cat .command.err - Read the task’s output log:
cat .command.out - See the exact command that ran:
cat .command.sh
The .command.sh file shows the exact command Nextflow ran. You can copy it and try running it yourself in an interactive job to debug.
2.9 Pipeline says it succeeded but results look wrong
Things to check:
- Open
multiqc/star_salmon/multiqc_report.htmland check alignment rates, very low alignment can produce a “successful” run with no useful data - Confirm the GTF annotation file matches the genome FASTA (both should come from the same EGAPx assembly)
- Check
pipeline_info/params_*.jsonto confirm the correct genome and annotation were used
3 Where to get more help
| Resource | Link |
|---|---|
| nf-core/rnaseq documentation | https://nf-co.re/rnaseq |
| nf-core Slack community | https://nf-co.re/join/slack |
| Unity documentation | https://docs.unity.rc.umass.edu |
| Unity support | Email hpc@umass.edu |