Snakemake workflow that sweeps over parameter grids, runs simulations in parallel, and aggregates results into a flat CSV.
workflow/
├── Snakefile # top-level entrypoint (run from project root)
├── rules/ # *.smk rule definitions
│ ├── simulate.smk
│ └── aggregate.smk
├── scripts/ # per-rule Python jobs
│ ├── simulate.py
│ └── aggregate.py
└── tools/ # shared helpers
├── sweeper.py # Config / ConfigSet — parameter-grid generation
└── job_utils.py # seed derivation, job/run metadata capture
experiments/<stage>/<exp>/sweeps/ # generated parameter grids and summary CSVs
A stage is a named simulation family (e.g. simulation, monte_carlo) with
its own parameter grid, sweep CSV, and output namespace.
A step is one computation phase within a stage.
The two built-in steps:
| Step | Rule | Output |
|---|---|---|
simulate |
simulate |
results.json per (path, instance) |
aggregate |
aggregate |
results.csv per stage |
Active stage and steps are selected in workflow.yaml:
active:
stage: simulation
steps: [simulate, aggregate]Valid step combinations (later steps depend on earlier ones):
[simulate][simulate, aggregate]
Each stage has a sweep script under experiments/<stage>/<exp>/sweeps/sweep.py that
defines a param_grid dict and calls ConfigSet.generate_and_save to
materialise a summary.csv. Nested dicts flatten to dot-separated CSV columns.
Supported value types in param_grid:
| Value | Behaviour |
|---|---|
| scalar | broadcast to every combination |
list / range |
Cartesian product axis |
lambda p: scalar |
derived lazily from already-resolved params |
nested dict |
keys appear as parent.child columns |
Generate or regenerate the CSV after any edit:
python experiments/simulation/e0/sweeps/sweep.pyThe Snakefile expands targets over every short_path row × {1..num_instances}.
Each parameter point (one row in summary.csv) is run num_instances times.
The {instance} wildcard is 1, 2, … num_instances. Every job derives its
RNG seed deterministically from (path, instance) via SHA-256:
- Same job → same result.
- Different instances of the same point → different random draws.
--rerun-incompletereproduces failed instances exactly.
The seed is stored in results.json and carried into results.csv.
# Dry-run: show what would execute.
uv run python -m snakemake --snakefile workflow/Snakefile --cores 4 -n
# Real run, 4 local workers.
uv run python -m snakemake --snakefile workflow/Snakefile --cores 4
# Force-rerun everything (e.g. after changing a script).
uv run python -m snakemake --snakefile workflow/Snakefile --cores 4 --forceall
# Rerun only incomplete / failed jobs.
uv run python -m snakemake --snakefile workflow/Snakefile --cores 4 --rerun-incompleteSet container in workflow.yaml to a local Apptainer image (.sif) and
launch with --use-singularity:
# workflow.yaml
container: containers/simulation-project-template-0.1.0.sifuv run python -m snakemake --snakefile workflow/Snakefile --cores 4 --use-singularityThe host environment provides snakemake; each job runs inside the container
via the workflow's global container: directive.
See docs/hpc-workflow.md for the full cluster guide.
One-liner for reference:
uv run python -m snakemake --snakefile workflow/Snakefile \
--executor slurm --jobs 500 \
--default-resources slurm_account=<acct> slurm_partition=<part> \
--use-singularityOr use the lifecycle manager:
python hpc/lifecycle.py submit mycluster myproject/workflowexperiments/<stage>/<exp>/data/<short_path>/<instance>/
├── results.json # params + seed + simulation output
├── numbers.json # intermediate data written by simulate.py
├── simulate.log # per-job stdout/stderr
└── benchmark.txt # Snakemake per-rule timing
experiments/<stage>/<exp>/results/
└── results.csv # flat per-run rows: params + seed + simulation output
What happens at parse time (before any job runs):
configfileloaded →ACTIVE_STAGE,ACTIVE_STEPS,EXP,CSV.summary.csvread into_df— must exist or Snakemake aborts immediately.params_by_pathdict built from CSV rows, keyed byshort_path.instances = [1..num_instances].include:rule files —EXP,_df,params_by_path,instances,configare module-level globals available in every.smkfile. Do not shadow them.rule alltargets computed by_all_targets()fromACTIVE_STEPS.
Same stage, different parameter ranges — copy the sweep script:
cp -r experiments/simulation/e0/sweeps experiments/simulation/e1/sweeps
# edit e1/sweeps/sweep.py: update root to "experiments/simulation/e1", adjust param_grid
python experiments/simulation/e1/sweeps/sweep.pyThen point workflow.yaml at it:
stages:
simulation:
exp: e1
csv: experiments/simulation/e1/sweeps/summary.csvOld e0 outputs are untouched.
A per-(path, instance) computation that consumes data produced by earlier
steps. docs/templates/new_step/ ships the skeleton. Steps:
- Implement the computation logic in
src/<package>/and test it. - Copy
docs/templates/new_step/rules/compute_foo.smk→workflow/rules/compute_foo.smk; wire inputs/outputs to match your step. - Copy
docs/templates/new_step/scripts/compute_foo.py→workflow/scripts/compute_foo.py; call your library function. - Apply
Snakefile.additions.pyandworkflow.yaml.additions(oneinclude:, one_all_targets()branch, one entry inactive.steps).
An independent simulation family with its own CSV and output tree.
docs/templates/new_stage/ ships a ready-made skeleton: sweep script, rule,
script, and the two *.additions files.
- Copy
docs/templates/new_stage/→workflow/rules/simulate_bar.smk,workflow/scripts/simulate_bar.py,experiments/bar/e0/sweeps/sweep.py. - Edit
param_gridand replacebarwith your stage name throughout. - Run the sweep script, then register the stage in
workflow.yaml. - Apply
Snakefile.additions.py(oneinclude:, one stage branch in_all_targets()).