Run the Pipeline

On Unity, the pipeline runs as a SLURM batch job, not interactively in the terminal. You write a shell script, submit it with sbatch, and SLURM runs it on a compute node while you can log out and do other things.

1 Understanding the job script

A job script is a plain text file. It has two parts:

  1. #SBATCH lines at the top, these tell SLURM how much time and memory to reserve, and where to write the output
  2. Shell commands below, these are the actual commands that run

Here is what each #SBATCH line means:

#SBATCH --job-name=nfcore-rnaseq-test   # A label shown in the queue
#SBATCH --partition=cpu                  # Use CPU nodes (not GPU)
#SBATCH --cpus-per-task=8               # Reserve 8 CPU cores for the manager process
#SBATCH --mem=48G                        # Reserve 48 GB of RAM
#SBATCH --time=24:00:00                  # Maximum run time (hours:minutes:seconds)
#SBATCH --output=.../%j.out             # File for standard output (%j = job ID)
#SBATCH --error=.../%j.err              # File for error output
#SBATCH --mail-type=BEGIN,END,FAIL      # Email you when the job starts, ends, or fails
#SBATCH --mail-user=you@email.edu       # Your email address
Note

The --cpus-per-task=8 and --mem=48G here are for the Nextflow manager process, not the individual pipeline steps. The pipeline itself will request additional resources for each step through SLURM automatically.

2 Understanding the environment setup

Inside the script, several lines set up the environment before running Nextflow:

module purge                         # Clear any previously loaded modules
module load nextflow/26.04.1         # Load Nextflow
module load apptainer/latest         # Load Apptainer (the container system)

The export lines tell Apptainer and Nextflow where to store container images and temporary files:

export APPTAINER_CACHEDIR="..."      # Where Apptainer stores downloaded container images
export APPTAINER_TMPDIR="..."        # Temporary files during container builds
export NXF_APPTAINER_CACHEDIR="..."  # Where Nextflow looks for Apptainer images
export NXF_OPTS='-Xms1g -Xmx4g'    # Memory limits for the Nextflow Java process
export PROOT_NO_SECCOMP=1           # Required for Apptainer to work correctly on Unity

These variables must point to your scratch workspace so they do not fill up your home directory.

3 Step 1: Run a test first

Before running with your own data, run the built-in test profile. This uses a tiny dataset that comes with the pipeline and confirms everything is working correctly on Unity.

3.1 Create the script

On Unity, open a new file with nano:

nano ~/rnaseq_nf_core/nfcore-rnaseq-test.sh

Paste in the following. Replace YOUR_USERNAME with your Unity username and YOUR_WORKSPACE with your scratch workspace name (from ws_list):

#!/bin/bash
#SBATCH --job-name=nfcore-rnaseq-test
#SBATCH --partition=cpu
#SBATCH --cpus-per-task=8
#SBATCH --mem=48G
#SBATCH --time=24:00:00
#SBATCH --output=/home/YOUR_USERNAME/rnaseq_nf_core/job-logs/nfcore-rnaseq-test.%j.out
#SBATCH --error=/home/YOUR_USERNAME/rnaseq_nf_core/job-logs/nfcore-rnaseq-test.%j.err
#SBATCH --mail-type=BEGIN,END,FAIL
#SBATCH --mail-user=you@email.edu

set -euo pipefail

SCRATCH_RNASEQ=/scratch4/workspace/YOUR_USERNAME-YOUR_WORKSPACE

cd "$SCRATCH_RNASEQ"

module purge
module load nextflow/26.04.1
module load apptainer/latest

mkdir -p "$SCRATCH_RNASEQ/.apptainer/build-cache"
mkdir -p "$SCRATCH_RNASEQ/.apptainer/tmp"
mkdir -p "$SCRATCH_RNASEQ/.nextflow-apptainer-cache"

export APPTAINER_CACHEDIR="$SCRATCH_RNASEQ/.apptainer/build-cache"
export APPTAINER_TMPDIR="$SCRATCH_RNASEQ/.apptainer/tmp"
export NXF_APPTAINER_CACHEDIR="$SCRATCH_RNASEQ/.nextflow-apptainer-cache"
export NXF_OPTS='-Xms1g -Xmx4g'
export PROOT_NO_SECCOMP=1

# Remove any partial container downloads from a previous attempt
find "$NXF_APPTAINER_CACHEDIR" -type f -name "*.pulling.*" -delete || true

nextflow run nf-core/rnaseq \
  -r 3.26.0 \
  -profile test,unity \
  --outdir test_results \
  -resume

Save and close nano: Ctrl + O, Enter, Ctrl + X.

3.2 What the nextflow command means

nextflow run nf-core/rnaseq \   # Run the nf-core/rnaseq pipeline
  -r 3.26.0 \                   # Use version 3.26.0 (pin the version for reproducibility)
  -profile test,unity \          # Use the built-in test dataset AND Unity cluster settings
  --outdir test_results \        # Put all output in a folder called test_results
  -resume                        # Resume from cached steps if re-run

3.3 Submit and monitor the job

sbatch ~/rnaseq_nf_core/nfcore-rnaseq-test.sh

You will see something like:

Submitted batch job 12345678

Note the job ID, you will use it to check status or find your log files.

Check whether the job is running:

squeue --me

Watch the log file update in real time (replace 12345678 with your job ID):

tail -f ~/rnaseq_nf_core/job-logs/nfcore-rnaseq-test.12345678.out

Press Ctrl + C to stop watching.

Tip

The test profile uses a small built-in dataset. It can take a few hours because the first run needs to download Apptainer container images. Subsequent runs are faster because the containers are cached.

3.4 How to tell if the test succeeded

A successful test run ends with something like:

-[nf-core/rnaseq] Pipeline completed successfully -

If you see Pipeline completed with errors, check the .err log file for details.

4 Step 2: Run with your own samples

For a real run you need three things:

  1. A samplesheet CSV listing your samples and FASTQ file paths
  2. A genome FASTA and GTF annotation (see Frog Data Availability)
  3. A params JSON file that points to everything

4.1 Create the samplesheet

The samplesheet is a CSV file with one row per sample. Create it with nano:

nano /scratch4/workspace/YOUR_USERNAME-YOUR_WORKSPACE/samplesheet.csv

It must have exactly these four columns:

sample,fastq_1,fastq_2,strandedness
C45-1B,/work/pi_lmangiamele_smith_edu/03_26_flut_yale_rnaseq/C45-1B_S2_L002_R1_001.fastq.gz,/work/pi_lmangiamele_smith_edu/03_26_flut_yale_rnaseq/C45-1B_S2_L002_R2_001.fastq.gz,reverse
C45-1H,/work/pi_lmangiamele_smith_edu/03_26_flut_yale_rnaseq/C45-1H_S1_L002_R1_001.fastq.gz,/work/pi_lmangiamele_smith_edu/03_26_flut_yale_rnaseq/C45-1H_S1_L002_R2_001.fastq.gz,reverse

Each column:

Column What to put
sample A short name for the sample, no spaces
fastq_1 Full path to the R1 (forward) FASTQ file
fastq_2 Full path to the R2 (reverse) FASTQ file
strandedness reverse, forward, or unstranded (use reverse for most Illumina RNA kits)

4.2 Create the params file

Rather than putting all options on the command line, save them in a JSON file. Create it with nano:

nano /scratch4/workspace/YOUR_USERNAME-YOUR_WORKSPACE/params.json
{
  "input": "/scratch4/workspace/YOUR_USERNAME-YOUR_WORKSPACE/samplesheet.csv",
  "outdir": "/scratch4/workspace/YOUR_USERNAME-YOUR_WORKSPACE/results",
  "fasta": "/work/pi_lmangiamele_smith_edu/output_tadpole_plus_adult/complete.genomic.fna",
  "gtf": "/work/pi_lmangiamele_smith_edu/output_tadpole_plus_adult/complete.genomic.gtf",
  "igenomes_ignore": true,
  "aligner": "star_salmon",
  "pseudo_aligner": "salmon",
  "gtf_extra_attributes": "gene",
  "featurecounts_group_type": "transcript_biotype",
  "max_cpus": 48,
  "max_memory": "256.GB",
  "max_time": "72.h"
}

Replace YOUR_USERNAME and YOUR_WORKSPACE with your actual values.

4.3 Create the SLURM script

nano ~/rnaseq_nf_core/run-rnaseq.sh
#!/bin/bash
#SBATCH --job-name=nfcore-rnaseq
#SBATCH --partition=cpu
#SBATCH --cpus-per-task=8
#SBATCH --mem=48G
#SBATCH --time=3-00:00:00
#SBATCH --output=/home/YOUR_USERNAME/rnaseq_nf_core/job-logs/nfcore-rnaseq.%j.out
#SBATCH --error=/home/YOUR_USERNAME/rnaseq_nf_core/job-logs/nfcore-rnaseq.%j.err
#SBATCH --mail-type=BEGIN,END,FAIL
#SBATCH --mail-user=you@email.edu

set -euo pipefail

SCRATCH_RNASEQ=/scratch4/workspace/YOUR_USERNAME-YOUR_WORKSPACE
PARAMS="$SCRATCH_RNASEQ/params.json"

cd "$SCRATCH_RNASEQ"

module purge
module load nextflow/26.04.1
module load apptainer/latest

mkdir -p "$SCRATCH_RNASEQ/.apptainer/build-cache"
mkdir -p "$SCRATCH_RNASEQ/.apptainer/tmp"

export APPTAINER_CACHEDIR="$SCRATCH_RNASEQ/.apptainer/build-cache"
export APPTAINER_TMPDIR="$SCRATCH_RNASEQ/.apptainer/tmp"
export PROOT_TMP_DIR="$SCRATCH_RNASEQ/.apptainer/tmp"
export TMPDIR="$SCRATCH_RNASEQ/.apptainer/tmp"
export NXF_APPTAINER_CACHEDIR="$SCRATCH_RNASEQ/.nextflow-apptainer-cache"
export NXF_OPTS='-Xms1g -Xmx4g'
export PROOT_NO_SECCOMP=1

nextflow run nf-core/rnaseq \
  -r 3.26.0 \
  -profile unity \
  -params-file "$PARAMS" \
  -resume

4.4 Submit

sbatch ~/rnaseq_nf_core/run-rnaseq.sh

5 Checking on a running job

Task Command
See your jobs in the queue squeue --me
Watch the output log live tail -f ~/rnaseq_nf_core/job-logs/nfcore-rnaseq.JOBID.out
Watch the error log live tail -f ~/rnaseq_nf_core/job-logs/nfcore-rnaseq.JOBID.err
Cancel a job scancel JOBID

Replace JOBID with the number that sbatch printed when you submitted.

6 Resuming an interrupted run

If your job ran out of time or failed partway through, fix the problem and resubmit with the same script. The -resume flag is already included, Nextflow will skip steps that finished successfully and continue from where it stopped.

Warning

Only resume using the same scratch directory and the same params file. Changing the output path or key parameters invalidates the cached steps and forces a full restart.

7 Record keeping

Keep a copy of your params file and SLURM script alongside your final results. The pipeline_info/ folder inside the results directory also stores exact parameters, software versions, and execution logs automatically.