Using Unity

1 Unity resources

Bookmark these links, you will use them regularly.

Resource Link
Connecting to Unity via SSH https://docs.unityhpc.org/documentation/connecting/ssh/
Scratch workspaces https://docs.unity.rc.umass.edu/documentation/managing-files/hpc-workspace/
Module explorer (search available software) https://ood.unity.rc.umass.edu/pun/sys/module-explorer
Node specifications https://docs.unityhpc.org/documentation/cluster_specs/nodes/

2 Your username

Throughout this tutorial you will see prompts and paths that contain kfloer_smith_edu. That is Kate’s Unity username. Replace it with your own Unity username everywhere you see it.

For example, if your Unity username is jsmith_smith_edu, then:

/home/kfloer_smith_edu/          →   /home/jsmith_smith_edu/
/scratch4/workspace/kfloer_smith_edu-demo/   →   /scratch4/workspace/jsmith_smith_edu-demo/

If you are not sure what your username is, run:

whoami

3 Login nodes

When you first connect to Unity with SSH, you land on a login node.

Your prompt may look like this:

(base) kfloer_smith_edu@login1:~$

Login nodes are for light work, such as:

  • Navigating directories
  • Editing small text files
  • Checking file names
  • Submitting jobs
  • Looking at small log files

Do not run large RNA-seq pipelines directly on a login node.

Warning

If a command will use lots of CPU, memory, or disk for more than a few seconds, it probably should not run on the login node.

4 Interactive jobs

An interactive job gives you a temporary compute session where you can run commands directly.

Interactive jobs are useful for:

  • Testing commands
  • Running small demos
  • Debugging a pipeline
  • Checking that software loads correctly
  • Trying a short nf-core test run

In an interactive job, you still type commands yourself, but the commands run on a compute node instead of the login node.

srun --partition=cpu --cpus-per-task=4 --mem=16G --time=2:00:00 --pty bash

Example prompt after starting an interactive job:

(base) kfloer_smith_edu@gpu-node:~$
Note

Interactive jobs are good for learning and testing. They are not always the best choice for long production runs.

5 Compute jobs

A compute job is a job submitted to the cluster scheduler. Instead of typing commands one by one, you write a job script and submit it.

Compute jobs are useful for:

  • Full RNA-seq pipeline runs
  • Long analyses
  • Analyses that need many CPUs
  • Analyses that need lots of memory
  • Work that should keep running after you log out

Submit a job script with sbatch:

sbatch my-script.sh

Check on your jobs:

squeue --me

6 Which one should I use?

Task Where to run it
Log in to Unity Login node
Look at files with ls Login node
Edit samplesheet.csv Login node
Check available workspaces with ws_list Login node
Test a small command Interactive job
Run an nf-core test profile Interactive job or compute job
Run a full RNA-seq pipeline Compute job
Run a long analysis overnight Compute job

7 Project architecture

Before you start running analyses, think about where your files will live. Unity has three main storage areas with different purposes.

Location Path pattern Purpose Space Temporary?
Home directory /home/kfloer_smith_edu/ Scripts, config files, small notes Limited (~50 GB) No
Lab shared storage /work/pi_lmangiamele_smith_edu/ Shared data, reference genomes, final results Large No
Scratch workspace /scratch4/workspace/kfloer_smith_edu-WORKSPACE/ Pipeline runs, intermediate files Very large Yes: 30 days

A typical project setup looks like this:

/home/kfloer_smith_edu/
  rnaseq_nf_core/
    job-logs/          ← SLURM output and error files
    scripts/           ← Your .sh job scripts

/work/pi_lmangiamele_smith_edu/
  03_26_flut_yale_rnaseq/     ← Raw FASTQ files (input data)
  output_tadpole_plus_adult/  ← Reference genome and annotation
  results_final/              ← Final results you want to keep

/scratch4/workspace/kfloer_smith_edu-rnaseq/
  samplesheets/        ← Input samplesheet CSV
  results/             ← nf-core pipeline output (large, temporary)
  .nextflow-apptainer-cache/  ← Container image cache
Warning

Scratch workspaces expire after 30 days. Copy anything you want to keep (final results, count tables, MultiQC reports) to /work/pi_lmangiamele_smith_edu/ before the workspace expires.

The scratch workspace is where pipelines run because it is fast and has large capacity. Your home directory and the lab’s /work/ directory are for permanent storage.

8 Conda and mamba environments

Conda is a package manager that lets you install software and keep different versions of tools isolated from each other. Mamba is a faster drop-in replacement for conda, they use the same commands, mamba just solves environments more quickly.

An environment is a self-contained collection of software. You create one environment per project (or per tool) so that different tools with conflicting dependencies do not interfere with each other.

8.1 Check what conda/mamba is available

module load miniconda/latest
conda --version

Or if mamba is available:

mamba --version

8.2 Create a new environment

mamba create -n rnaseq-env python=3.11

8.3 Activate an environment

conda activate rnaseq-env

Your prompt will change to show the environment name:

(rnaseq-env) kfloer_smith_edu@login1:~$

8.4 Install packages into the active environment

mamba install -c bioconda -c conda-forge fastqc multiqc

8.5 Deactivate

conda deactivate

8.6 List your environments

conda env list
Note

For running nf-core pipelines on Unity, you generally do not need a conda environment, the pipeline manages its own software through Apptainer containers. Conda environments are more useful for tools you run separately, like R packages or custom Python scripts.

9 A useful mental model

Think of the login node as the front desk. It helps you get organized and submit work.

Think of interactive jobs as a temporary workbench. They are useful when you need to test something hands-on.

Think of compute jobs as the production workspace. They are where long or heavy analyses should run.