Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -14,7 +14,7 @@ jobs:
working-directory: cli_tools/mcdi
strategy:
matrix:
python-version: ["3.9", "3.11"]
python-version: ["3.10", "3.12"]
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
Expand Down
2 changes: 1 addition & 1 deletion cli_tools/mcdi/Dockerfile
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
FROM python:3.12-alpine
FROM python:3.12-slim
WORKDIR /app
# build from the mcdi parent directory
COPY . .
Expand Down
103 changes: 52 additions & 51 deletions cli_tools/mcdi/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,18 +4,18 @@ Galaxy Cancer Data Importers (GaCDI) provides Galaxy tools for importing cancer
datasets from major public and controlled-access cancer data repositories into
Galaxy histories. This package provides one command, `mcdi` (Multi-Commons Data
Importer), with two subcommands: `mcdi manifest` builds a manifest from
filters, and `mcdi download` downloads the files a GDC or PDC manifest lists
(whether built here or exported from a portal).
filters, and `mcdi download` downloads the files a GDC, PDC, or IDC manifest
lists (whether built here or exported from a portal).

## Manifest Builder (this branch)
## Manifest Builder

`gacdi_manifest_gdc` generates the **manifests** that drive the importers. Instead
of downloading a whole dataset, the user filters the NCI
[Genomic Data Commons](https://gdc.cancer.gov/) and gets exactly the files they
want, described in two complementary outputs:

- **GDC manifest** — strict `id / filename / md5 / size / state`, consumable
directly by `gdc-client` and by the GaCDI GDC importer.
directly by `gdc-client` and by `mcdi download` (the GaCDI Manifest Downloader).
- **metadata table** — the same files joined to sample barcodes and, optionally,
clinical/molecular annotations (GDC fields, cBioPortal subtypes like PAM50 and
ER/PR/HER2, and/or a user-uploaded annotation TSV). It also carries
Expand Down Expand Up @@ -58,40 +58,34 @@ repeatable custom facets (`--extra-filter "field=…;op=in|exclude;values=a,b"`)
a raw GDC filters JSON (`--raw-filters`). The manifest is emitted in a deterministic
(sorted) order for reproducible workflows.

### End-to-end with the GaCDI GDC importer
### Feeding the manifest into `mcdi download`

This tool is designed to feed directly into the **GaCDI GDC importer** (the
manifest-download branch), so a single Galaxy workflow goes *filter → manifest →
download → analysis*:
A single Galaxy workflow goes *filter → manifest → download → analysis*: build
the manifest here, run **GaCDI Manifest Downloader** (`mcdi download`, below)
to bring the files into the history, then join `metadata.tsv` to those history
datasets to attach clinical labels, the `galaxy_ext` datatype hint, and subtype
annotations to each sample.

```
[GaCDI Manifest Builder]--gdc_manifest.txt-->[GaCDI GDC importer]--collection-->[analysis tools]
\--metadata.tsv------------------------------(join)----/
```

**Compatibility contract (locked by `tests/test_importer_contract.py`):**
This works because of a contract between the two tools (locked by
`tests/test_importer_contract.py`):

1. **Manifest → importer.** `gdc_manifest.txt` is a TSV whose header
(`id, filename, md5, size, state`) is a superset of what the importer's
`parse_gdc_manifest` requires (`id/filename/md5/size`); its datatype (`txt`)
is accepted by the importer's manifest input (`tabular,txt`). Rows with no
`id` are dropped so the manifest and metadata stay row-aligned. The same file
also works with `gdc-client download -m gdc_manifest.txt`.
1. **Manifest → downloader.** `gdc_manifest.txt` is a TSV whose header
(`id, filename, md5, size, state`) is a superset of what `mcdi download`'s
GDC parsing requires (`id/filename/md5/size`); its datatype (`txt`) is
accepted by the downloader's manifest input (`tabular,txt`). Rows with no
`id` are dropped so the manifest and metadata stay row-aligned. The same
file also works with `gdc-client download -m gdc_manifest.txt`.
2. **Metadata ↔ history.** `metadata.tsv` leads with `file_id` and `filename` —
the exact keys of the importer's history **summary** — so after download you
join the metadata to the imported collection/summary on `file_id` (stable
UUID) or `filename` to attach clinical labels, the `galaxy_ext` datatype hint,
and subtype annotations to each sample in the history.

So: build the manifest here, run the importer to bring the samples into the
history, then join `metadata.tsv` to give those history datasets their
annotations (e.g. labels for an image ML model).
the same keys the downloaded collection's datasets are named/identified by
— so after download you join the metadata to the collection on `file_id`
(stable UUID) or `filename` to attach clinical labels, the `galaxy_ext`
datatype hint, and subtype annotations to each sample in the history.

## Downloading files from a manifest

`mcdi download` fetches the files listed in a GDC or PDC manifest — either
one built by `mcdi manifest gdc` above, or one exported directly from a
portal:
`mcdi download` fetches the files listed in a GDC, PDC, or IDC manifest —
either one built by `mcdi manifest gdc` above, or one exported directly from
a portal:

- **GDC**: build a file cart in the [GDC portal](https://portal.gdc.cancer.gov)
and download the manifest (TSV), or generate one via the API:
Expand All @@ -100,22 +94,32 @@ portal:
to the files you want, and use "Export File Manifest" (CSV or TSV). Note that
the signed download links embedded in a PDC manifest expire after 7 days —
re-export if downloads start failing.
- **IDC**: build a cohort in the [IDC portal](https://portal.imaging.datacommons.cancer.gov)
and export either manifest format it offers — the **s5cmd manifest** (a
`.s5cmd` script of `cp s3://...` lines, one per DICOM series) or the
**CSV/TSV cohort manifest** (a table keyed by `SeriesInstanceUID`). Both are
auto-detected and accepted. A manifest entry references a whole series (many
files, not one), so the actual file list is resolved at download time
against [idc-index](https://github.com/ImagingDataCommons/idc-index)'s
metadata and the public `idc-open-data` S3 bucket; only AWS-hosted series are
currently supported. No token is needed — all IDC data is open access.

```bash
mcdi download --manifest gdc_manifest.txt --output-dir downloads/
mcdi download --manifest pdc_manifest.csv --output-dir downloads/ --verify-checksum
mcdi download --manifest idc_manifest.s5cmd --output-dir downloads/
```

The data commons is auto-detected from the manifest's header row; pass
`--source {gdc,pdc}` to override.
The data commons is auto-detected from the manifest's content; pass
`--source {gdc,idc,pdc}` to override.

| Flag | Description |
|---|---|
| `--manifest PATH` | Path to the manifest (required) |
| `--output-dir DIR` | Directory to download files into (required) |
| `--source {gdc,pdc}` | Skip auto-detection |
| `--source {gdc,idc,pdc}` | Skip auto-detection |
| `--workers N` | Concurrent downloads (default: 4). PDC always runs at 1 to respect its rate limit, regardless of this flag. |
| `--verify-checksum` | Verify each file's md5 against the manifest after download |
| `--verify-checksum` | Verify each file's md5 against the manifest after download. No-op for IDC entries: IDC has no reliable per-object checksum (S3's ETag isn't sound for multipart uploads), so `md5` is always unset and verification is silently skipped, same as any manifest entry without one. |
| `-x`, `--extract` | Extract recognized archives (`.tar.gz`, `.tgz`, `.tar.bz2`, `.tar.xz`, `.tar`, `.zip`, `.gz`, `.bz2`, `.xz`) in place after download, keeping only the extracted contents at the manifest's output path. Off by default. |
| `--token-file PATH` | File containing a GDC auth token, for controlled-access files |
| `--retries N` | Extra attempts for files that fail transiently within this run (default: 2) |
Expand Down Expand Up @@ -163,11 +167,14 @@ mcdi download --manifest gdc_manifest.txt --output-dir downloads/
```

PDC downloads use pre-signed URLs embedded in the manifest and need no token.
IDC downloads are open access and also need no token.

Output layout:

- **GDC**: `downloads/gdc/<file_id>/<filename>`
- **PDC**: `downloads/pdc/<study_id>/<study_version>/<data_category>/[<run_metadata_id>/]<file_type>/<filename>`
- **IDC**: `downloads/idc/<collection_id>/<PatientID>/<StudyInstanceUID>/<Modality>_<SeriesInstanceUID>/<filename>`
(mirrors idc-index's own default download hierarchy)

Re-running against the same manifest and output directory skips files already
downloaded (verifying checksums too, if `--verify-checksum` is set), so
Expand All @@ -176,33 +183,27 @@ so a rerun over mostly-complete output is cheap, not another full pass.

## Runtime environment

The tool ships a pinned container (`quay.io/<org>/mcdi`) referenced from
the wrapper, with Python + `requests` Conda requirements as a fallback. The Quay
namespace (`paulocilasjr`) is a placeholder — update `@QUAY_ORG@` in
`tools/manifest_gdc/macros.xml`, `containers/Dockerfile.manifest`, and the workflow
before publishing.
Both Galaxy tools (`manifest_gdc`, `manifest_downloader`) reference the same
pinned container, `quay.io/goeckslab/mcdi:<version>` — there's no Conda
package, so the container is the sole runtime. It's built from
`cli_tools/mcdi/Dockerfile`, tagged from `mcdi/__init__.py`'s `__version__`,
and built/pushed automatically by `.github/workflows/containers.yml` on
merges to `main` that touch `cli_tools/mcdi/**`.

```bash
docker build -f containers/Dockerfile.manifest -t mcdi:dev .
docker build -t mcdi:dev cli_tools/mcdi
docker run --rm mcdi:dev mcdi manifest gdc --help
docker run --rm mcdi:dev mcdi download --help
```

## Development

```bash
python -m pip install -e '.[dev]'
pytest -q # mocked; no network
planemo lint tools/manifest_gdc
pytest -q -m "not network" # mocked; add `-m network` for live-API tests
planemo lint tools/manifest_gdc tools/manifest_downloader
```

## Roadmap

- **Phase 1 (this branch):** GDC manifest builder + enrichment + join/QC.
- **Phase 2:** CRDC GDC-style commons (PDC/IDC/ICDC/CDS/CTDC) reusing the filter/
join core.
- **Phase 3:** GEO/SRA accession-list builders; on merge with the importer branch,
fold shared HTTP utilities into the `gacdi` package.

## License

See [LICENSE](LICENSE).
19 changes: 4 additions & 15 deletions cli_tools/mcdi/mcdi/__init__.py
Original file line number Diff line number Diff line change
@@ -1,25 +1,14 @@
"""MCDI — Multi-Commons Data Importer.

One command, two subcommands:

- ``mcdi manifest`` (:mod:`mcdi.manifest`): filter-driven generation of
download manifests (and enriched metadata tables) for NIH/NCI cancer data
repositories, starting with the NCI Genomic Data Commons (GDC). The builder
emits a deliberate *two-file split*: a lean, CLI/importer-ready manifest
(``id/filename/md5/size/state``) and a rich metadata table joining
clinical/molecular annotations by barcode, plus a match report so selections
and joins are never silently wrong.
- ``mcdi download`` (:mod:`mcdi.download`): downloads the files listed
in a GDC or PDC manifest, auto-detecting which commons it came from.
``mcdi manifest``: build a filtered GDC manifest + enriched metadata table.
``mcdi download``: download files from a GDC, PDC, or IDC manifest.
"""

import os

__version__ = "0.4.0"
__version__ = "0.5.0"

# Build identifier baked into the container image at build time (e.g. the git
# commit SHA). Lets you confirm the exact code a run used, even when the version
# number hasn't changed. Empty for local/editable installs.
# Container image build id (e.g. git SHA); empty for local/editable installs.
BUILD = os.environ.get("MCDI_BUILD", "").strip()


Expand Down
4 changes: 2 additions & 2 deletions cli_tools/mcdi/mcdi/download/cli.py
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
"""``mcdi download`` — download the files listed in a GDC or PDC manifest."""
"""``mcdi download`` — download the files listed in a GDC, PDC, or IDC manifest."""

from __future__ import annotations

Expand All @@ -17,7 +17,7 @@ def add_arguments(subparsers: argparse._SubParsersAction) -> argparse.ArgumentPa
"""Attach the ``download`` subcommand to ``subparsers``."""
parser = subparsers.add_parser(
"download",
help="Download files listed in a GDC or PDC manifest.",
help="Download files listed in a GDC, PDC, or IDC manifest.",
)
parser.add_argument("--manifest", required=True, type=Path, help="Path to the exported manifest file")
parser.add_argument("--output-dir", required=True, type=Path, help="Directory to download files into")
Expand Down
53 changes: 13 additions & 40 deletions cli_tools/mcdi/mcdi/download/engine.py
Original file line number Diff line number Diff line change
Expand Up @@ -71,9 +71,8 @@ def _dest_path(output_dir: Path, entry: FileEntry) -> Path:
def _already_present(dest_path: Path, entry: FileEntry, verify: bool) -> tuple[bool, str]:
"""Return ``(satisfied, detail)``: whether ``dest_path`` already correctly holds ``entry``.

Shared by the download step and the pre-flight access check, so a file
that doesn't need (re-)downloading also doesn't need its remote
accessibility re-verified on every rerun.
Shared by the download step and the pre-flight access check, so a rerun
skips both re-downloading and re-verifying access for files already present.
"""
if not (dest_path.exists() and dest_path.stat().st_size > 0):
return False, ""
Expand All @@ -94,19 +93,10 @@ class DownloadResult:


def _archived_path(output_dir: Path, entry: FileEntry) -> Path:
"""Where ``entry``'s archive is relocated to once successfully extracted.

Mirrors ``entry.rel_dir``/filename under a sibling of ``output_dir``
(``<output_dir>.mcdi-archives/...``). Moving the archive there - instead
of leaving it in ``output_dir`` next to what it was extracted into -
means (a) a tool that recursively collects everything under
``output_dir`` (e.g. a Galaxy ``discover_datasets`` with ``recurse``)
only ever sees the extracted contents, not a redundant copy of the
packed archive, and (b) the archive's presence here doubles as the
idempotency marker: extraction is already done for this entry iff a file
exists here. The archive isn't deleted, just moved aside - if extraction
fails, it's left in ``output_dir`` instead, so there's still something
to show for the download.
"""Where ``entry``'s archive moves to after successful extraction: a sibling
``<output_dir>.mcdi-archives/`` tree, mirroring ``entry.rel_dir``.

Its presence there doubles as the idempotency marker for "already extracted".
"""
archive_root = output_dir.with_name(output_dir.name + ".mcdi-archives")
return archive_root / entry.rel_dir / entry.filename
Expand Down Expand Up @@ -187,15 +177,12 @@ def _download_one(
return result


# HTTP statuses worth retrying at the batch level: rate limiting and
# server-side/transient failures. Anything else (401/403/404/...) is a
# permanent-looking failure that a delayed retry won't fix.
# Transient: rate limiting + server errors. Not 401/403/404 (permanent).
_RETRYABLE_STATUSES = {429, 500, 502, 503, 504}


def _is_retryable(result: DownloadResult) -> bool:
# A checksum mismatch could be transient transfer corruption, not
# necessarily a bad source file, so it's worth one more attempt too.
# A checksum mismatch could be transient transfer corruption, not a bad source file.
if result.status == "checksum_mismatch":
return True
if result.status != "error":
Expand All @@ -206,8 +193,7 @@ def _is_retryable(result: DownloadResult) -> bool:
return int(detail.split()[1]) in _RETRYABLE_STATUSES
except (IndexError, ValueError):
return False
# Any other "error" detail came from a requests.RequestException (timeout,
# connection reset, DNS hiccup, ...) - inherently transient.
# Any other "error" detail is a requests.RequestException (timeout, connection reset, ...) - transient.
return True


Expand All @@ -224,8 +210,7 @@ def run(
session = build_session()
rate_limit = source.rate_limit()
pacer = Pacer(rate_limit) if rate_limit else None
# PDC's per-IP rate limit applies regardless of thread count, so cap
# effective concurrency to 1 when a pacer is active to keep pacing honest.
# Cap concurrency to 1 when paced (e.g. PDC's per-IP limit), regardless of --workers.
effective_workers = 1 if pacer else workers

def _pass(batch: list[FileEntry]) -> list[DownloadResult]:
Expand All @@ -248,10 +233,7 @@ def _pass(batch: list[FileEntry]) -> list[DownloadResult]:

results_by_id = {r.entry.file_id: r for r in _pass(entries)}

# Batch-level retry, on top of the per-request transport retry already
# inside `build_session()`. This matters most for non-interactive runs
# (e.g. a Galaxy job) where nothing will manually rerun the command on
# the same output directory if a few files fail transiently.
# Batch-level retry, on top of build_session()'s own transport-level retries.
attempt = 0
while attempt < retries:
retry_entries = [r.entry for r in results_by_id.values() if _is_retryable(r)]
Expand Down Expand Up @@ -302,17 +284,8 @@ def check_access(
) -> list[AccessFailure]:
"""Probe every entry with a 1-byte ranged request; return the ones that fail.

Meant to run before any real download, so a manifest with one
inaccessible file (e.g. controlled-access without a valid token) is
caught in seconds instead of after downloading everything else first.
Two things narrow what actually needs a round trip: entries already
correctly present in ``output_dir`` (same check the download step uses -
including, if ``extract`` is set, entries already extracted and relocated
to ``output_dir``'s sibling archive directory - so a rerun over
mostly-complete output doesn't re-verify remote access for files it isn't
going to touch anyway) skip the check outright, and entries the source
can already vouch for as open (see ``Source.known_open``) skip the
per-file probe specifically.
Skips entries already present in ``output_dir`` (or already extracted, if
``extract``) and ones ``Source.known_open`` vouches for.
"""
session = build_session()
rate_limit = source.rate_limit()
Expand Down
14 changes: 10 additions & 4 deletions cli_tools/mcdi/mcdi/download/sources/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -3,24 +3,28 @@
from pathlib import Path

from ...errors import InputError
from .base import CANDIDATE_DELIMITERS, FileEntry, RateLimit, Source, read_header
from .base import CANDIDATE_DELIMITERS, FileEntry, RateLimit, Source, read_header, read_lines
from .gdc import GDCSource
from .idc import IDCSource
from .pdc import PDCSource

SOURCES: dict[str, type[Source]] = {
GDCSource.name: GDCSource,
PDCSource.name: PDCSource,
IDCSource.name: IDCSource,
}


def detect_source(path: Path) -> str:
"""Identify which data commons a manifest came from by matching its header
row, under each candidate delimiter, against every registered source's
schema (``Source.sniff``)."""
schema (``Source.sniff``), or - for headerless formats - its leading raw
lines (``Source.sniff_lines``)."""
lines = read_lines(path)
matches = [
name
for name, cls in SOURCES.items()
if any(cls.sniff(read_header(path, d)) for d in CANDIDATE_DELIMITERS)
if any(cls.sniff(read_header(path, d)) for d in CANDIDATE_DELIMITERS) or cls.sniff_lines(lines)
]
if len(matches) == 1:
return matches[0]
Expand All @@ -34,4 +38,6 @@ def detect_source(path: Path) -> str:
)


__all__ = ["FileEntry", "RateLimit", "Source", "GDCSource", "PDCSource", "SOURCES", "detect_source"]
__all__ = [
"FileEntry", "RateLimit", "Source", "GDCSource", "PDCSource", "IDCSource", "SOURCES", "detect_source",
]
Loading
Loading