Running the Snakemake pipeline on a SLURM cluster via Apptainer/Singularity.
Per-simulation knobs live in gitignored files (hpc/config/config.yaml,
hpc/config/submit.yaml), so the tracked code stays stable across experiments.
hpc/ ← tracked
├── lifecycle.py # CLI: upload / setup / submit / check / download / status
├── config/
│ ├── config_example.yaml # Snakemake SLURM profile template
│ ├── submit_example.yaml # sbatch launcher-job resources template
│ └── clusters_example.yaml # SSH host → remote root path mapping template
├── setup/
│ ├── setup.sh # one-time cluster venv bootstrap
│ └── requirements.txt # pinned Snakemake + executor-plugin deps
├── snakemake/
│ └── run_snakemake.sh # cluster job entrypoint (called by sbatch)
└── containers/
├── build_sif.py # Docker → OCI archive → Apptainer SIF
└── build_docker.sh # local Docker build + run (dev / testing)
hpc/config/config.yaml ← gitignored, copy from config_example.yaml
hpc/config/submit.yaml ← gitignored, copy from submit_example.yaml
hpc/config/clusters.yaml ← gitignored, copy from clusters_example.yaml
| Concern | File | Tracked? |
|---|---|---|
| HPC orchestration code | hpc/*.py, hpc/**/*.sh |
yes |
| Per-rule SLURM defaults (memory, cpus, runtime) | hpc/config/config_example.yaml |
yes (example) |
| Active Snakemake cluster profile | hpc/config/config.yaml |
no |
| Launcher job resources (wall time, mail) | hpc/config/submit_example.yaml |
yes (example) |
| Launcher job resources (your values) | hpc/config/submit.yaml |
no |
| SSH host → remote root mapping | hpc/config/clusters_example.yaml |
yes (example) |
| Cluster root paths (your values) | hpc/config/clusters.yaml |
no |
| Experiment parameters (stage, CSV, instances) | workflow.yaml |
yes |
| SIF / OCI image artifacts | containers/*.sif, *.tar.gz |
no |
To change resources or which experiment runs, only touch the gitignored files — the workflow code itself never needs editing per simulation.
From the project root:
# 1. Create config files from the example templates.
cp hpc/config/config_example.yaml hpc/config/config.yaml # adjust account, partition, resources
cp hpc/config/submit_example.yaml hpc/config/submit.yaml # adjust wall time, mail, account
cp hpc/config/clusters_example.yaml hpc/config/clusters.yaml # fill in your cluster paths
# 2. (Optional) Build the Apptainer SIF image locally.
uv run hpc/containers/build_sif.py --platforms linux/amd64
# 3. Generate the parameter-grid CSV.
python experiments/simulation/e0/sweeps/sweep.py
# 4. Upload workflow files (and the SIF if built) to the cluster.
python hpc/lifecycle.py upload mycluster myproject/workflow
# 5. One-time on the cluster: create the venv with Snakemake.
python hpc/lifecycle.py setup mycluster myproject/workflow
# If your site requires a compute allocation for pip installs:
python hpc/lifecycle.py setup mycluster myproject/workflow --sallocmycluster is an SSH host alias (defined in ~/.ssh/config).
myproject/workflow is the path under the cluster root where the project lands.
hpc/lifecycle.py reads cluster root paths from hpc/config/clusters.yaml
(gitignored). Copy the example and fill in your paths:
cp hpc/config/clusters_example.yaml hpc/config/clusters.yaml# hpc/config/clusters.yaml
clusters:
mycluster: /scratch/myuser
default_root: /home/myuser/projects/def-pi/myuserOr pass --remote-root /path/to/root on every command.
# Submit (profile mode: one SLURM job per rule instance -> good for long jobs).
python hpc/lifecycle.py submit mycluster myproject/workflow
# Submit (local mode: single job using all allocated CPUs -> good for quick jobs)
python hpc/lifecycle.py submit mycluster myproject/workflow --mode local
# Plan only — show what Snakemake would run without submitting.
python hpc/lifecycle.py submit mycluster myproject/workflow --snakemake-dry-run
# Monitor.
python hpc/lifecycle.py status mycluster # squeue
python hpc/lifecycle.py status mycluster --full # sacct history
python hpc/lifecycle.py check mycluster myproject/workflow # snakemake --summary
# Pull results back.
python hpc/lifecycle.py download mycluster myproject/workflow --paths results
python hpc/lifecycle.py download mycluster myproject/workflow --include-data # + data/local: hpc/config/submit.yaml ──┐
│ rsync via lifecycle.py upload
local: workflow.yaml ┤
local: experiments/*/e*/sweeps/summary.csv ┘
▼
cluster: <remote_folder>/ (same paths, read by lifecycle.py submit + run_snakemake.sh)
│
▼
sbatch <submit.yaml flags> hpc/snakemake/run_snakemake.sh <mode>
│
▼
snakemake --snakefile workflow/Snakefile --profile hpc/config ...
hpc/config/submit.yaml controls the launcher job that owns the Snakemake process.
hpc/config/config.yaml controls each per-rule job dispatched by Snakemake.
# Build + run with auto-mounted ./configs and ./results.
./hpc/containers/build_docker.sh
# Pass CLI args directly.
./hpc/containers/build_docker.sh generate 100 -o /results/out.json# linux/amd64 SIF (default).
uv run hpc/containers/build_sif.py
# Multi-platform.
uv run hpc/containers/build_sif.py --platforms linux/amd64 linux/arm64The SIF lands in containers/. Point workflow.yaml at it before uploading:
container: containers/simulation-project-template-0.1.0.sifhpc/setup/requirements.txt lists the Snakemake packages installed on the
cluster venv. Update it when bumping Snakemake:
# Pin only the packages needed on the cluster (no project package itself).
uv export --no-hashes --no-emit-project \
--only-package snakemake \
--only-package snakemake-executor-plugin-slurm \
--only-package importlib-metadata \
> hpc/setup/requirements.txt