Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
116 changes: 65 additions & 51 deletions cli_tools/mcdi/README.md
Original file line number Diff line number Diff line change
@@ -1,17 +1,41 @@
# GaCDI — Galaxy Cancer Data Importers
# mcdi — Multi-Commons Data Importer

Galaxy Cancer Data Importers (GaCDI) provides Galaxy tools for importing cancer
datasets from major public and controlled-access cancer data repositories into
Galaxy histories. This package provides one command, `mcdi` (Multi-Commons Data
Importer), with two subcommands: `mcdi manifest` builds a manifest from
filters, and `mcdi download` downloads the files a GDC, PDC, or IDC manifest
lists (whether built here or exported from a portal).
`mcdi` is a command-line tool for importing cancer datasets from major public
and controlled-access cancer data repositories: the NCI Genomic Data Commons
([GDC](https://portal.gdc.cancer.gov)), Proteomic Data Commons
([PDC](https://pdc.cancer.gov)), and Imaging Data Commons
([IDC](https://portal.imaging.datacommons.cancer.gov)). It has two
subcommands: `mcdi manifest` builds a manifest from filters, and `mcdi
download` downloads the files a GDC, PDC, or IDC manifest lists.

Manifests are usually exported directly from one of those three portals;
`mcdi download` accepts any of their native export formats as-is. `mcdi
manifest gdc` additionally builds a GDC manifest — plus enriched metadata —
straight from filters, as an alternative to the portal UI.

`mcdi` also powers the GaCDI (Galaxy Cancer Data Importers) Galaxy tools
published in this repo (`manifest_gdc`, `manifest_downloader`) — see their
own help text for Galaxy-specific usage.

## Installation

```bash
pip install cli_tools/mcdi
```

Or run the published container directly, without installing anything locally
(the same image also backs the repo's Galaxy tools):

```bash
docker run --rm quay.io/goeckslab/mcdi:<version> mcdi manifest gdc --help
docker run --rm quay.io/goeckslab/mcdi:<version> mcdi download --help
```

## Manifest Builder

`gacdi_manifest_gdc` generates the **manifests** that drive the importers. Instead
of downloading a whole dataset, the user filters the NCI
[Genomic Data Commons](https://gdc.cancer.gov/) and gets exactly the files they
`mcdi manifest gdc` builds the **manifests** that drive the downloader.
Instead of downloading a whole dataset, you filter the NCI
[Genomic Data Commons](https://gdc.cancer.gov/) and get exactly the files you
want, described in two complementary outputs:

- **GDC manifest** — strict `id / filename / md5 / size / state`, consumable
Expand Down Expand Up @@ -60,26 +84,24 @@ a raw GDC filters JSON (`--raw-filters`). The manifest is emitted in a determini

### Feeding the manifest into `mcdi download`

A single Galaxy workflow goes *filter → manifest → download → analysis*: build
the manifest here, run **GaCDI Manifest Downloader** (`mcdi download`, below)
to bring the files into the history, then join `metadata.tsv` to those history
datasets to attach clinical labels, the `galaxy_ext` datatype hint, and subtype
annotations to each sample.
The usual flow is *filter → manifest → download → analysis*: build the
manifest here, run `mcdi download` (below) to fetch the files, then join
`metadata.tsv` to the downloaded files to attach clinical labels, the
`galaxy_ext` datatype hint, and subtype annotations to each sample. (This
hand-off is also what powers the paired GaCDI Galaxy tools.)

This works because of a contract between the two tools (locked by
`tests/test_importer_contract.py`):
This works because of a contract between the two tools:

1. **Manifest → downloader.** `gdc_manifest.txt` is a TSV whose header
(`id, filename, md5, size, state`) is a superset of what `mcdi download`'s
GDC parsing requires (`id/filename/md5/size`); its datatype (`txt`) is
accepted by the downloader's manifest input (`tabular,txt`). Rows with no
`id` are dropped so the manifest and metadata stay row-aligned. The same
file also works with `gdc-client download -m gdc_manifest.txt`.
2. **Metadata ↔ history.** `metadata.tsv` leads with `file_id` and `filename` —
the same keys the downloaded collection's datasets are named/identified by
— so after download you join the metadata to the collection on `file_id`
(stable UUID) or `filename` to attach clinical labels, the `galaxy_ext`
datatype hint, and subtype annotations to each sample in the history.
GDC parsing requires (`id/filename/md5/size`). Rows with no `id` are
dropped so the manifest and metadata stay row-aligned. The same file also
works with `gdc-client download -m gdc_manifest.txt`.
2. **Metadata ↔ downloaded files.** `metadata.tsv` leads with `file_id` and
`filename` — the same identifiers `mcdi download` organizes its output by
(see "Output layout" below) — so after download you join the metadata on
`file_id` (stable UUID) or `filename` to attach clinical labels, the
`galaxy_ext` datatype hint, and subtype annotations to each sample.

## Downloading files from a manifest

Expand Down Expand Up @@ -124,16 +146,17 @@ The data commons is auto-detected from the manifest's content; pass
| `--token-file PATH` | File containing a GDC auth token, for controlled-access files |
| `--retries N` | Extra attempts for files that fail transiently within this run (default: 2) |
| `--retry-backoff SECONDS` | Wait before each retry pass, multiplied by the attempt number (default: 5.0) |
| `--retry-manifest-out PATH` | Where to write the retry manifest on a partial failure (default: `<manifest>.retry<ext>`, next to `--manifest`) |

Some commons files are themselves archives (e.g. a `.tar.gz` bundle of slides).
`--extract` unpacks any recognized archive into the same directory it was
downloaded into, right after downloading it. On success, the archive itself
then moves to a sibling `<output-dir>.mcdi-archives/` directory (mirroring
`--output-dir`'s layout) — it isn't deleted, just relocated out of the way,
so `--output-dir` ends up holding only the extracted contents, not a
redundant copy of the packed archive next to them. A tool that recursively
collects everything under `--output-dir` (e.g. Galaxy's `discover_datasets`)
then only ever sees the actual extracted files. If extraction fails, the
redundant copy of the packed archive next to them. Anything that
recursively scans `--output-dir` afterward only ever sees the actual
extracted files. If extraction fails, the
archive is left where it was downloaded instead, so there's still something
to show for it. The archive's presence in `.mcdi-archives/` also doubles as
the idempotency marker, so reruns skip both re-downloading and
Expand All @@ -152,10 +175,16 @@ file. This always runs; there's no flag to skip it.
request, a failed file (connection errors, `429`/`5xx`, or a checksum
mismatch — not permanent-looking failures like `401`/`403`/`404`) gets
`--retries` more whole-batch attempts, waiting `--retry-backoff × attempt`
seconds between passes. This matters most where nothing will manually rerun
the command for you on a failure — e.g. a Galaxy job, where a retried job
gets a fresh working directory, not the partial output of the failed attempt,
so anything not resolved within the one invocation is lost.
seconds between passes. This matters most in automated contexts where
nothing else will rerun the command for you on a failure.

**Partial success.** If some files succeed and others still fail after
retries, the run doesn't just fail outright — it also writes a manifest of
just the failed files (`<manifest>.retry<ext>`, or `--retry-manifest-out`'s
path if given) so you can rerun `mcdi download` against only those, instead
of the whole original manifest. The exit code tells you which case you're
in: `0` clean success, `3` partial success (some files missing, but at least
one succeeded), `1` total failure (nothing usable was produced).

Controlled-access GDC files need an auth token, obtained by logging into the
GDC portal and downloading your token. Provide it either via the `GDC_TOKEN`
Expand All @@ -181,29 +210,14 @@ downloaded (verifying checksums too, if `--verify-checksum` is set), so
interrupted runs can simply be re-run — the pre-flight check skips them too,
so a rerun over mostly-complete output is cheap, not another full pass.

## Runtime environment

Both Galaxy tools (`manifest_gdc`, `manifest_downloader`) reference the same
pinned container, `quay.io/goeckslab/mcdi:<version>` — there's no Conda
package, so the container is the sole runtime. It's built from
`cli_tools/mcdi/Dockerfile`, tagged from `mcdi/__init__.py`'s `__version__`,
and built/pushed automatically by `.github/workflows/containers.yml` on
merges to `main` that touch `cli_tools/mcdi/**`.
## License

```bash
docker build -t mcdi:dev cli_tools/mcdi
docker run --rm mcdi:dev mcdi manifest gdc --help
docker run --rm mcdi:dev mcdi download --help
```
See [LICENSE](LICENSE).

## Development
## Contributing

```bash
python -m pip install -e '.[dev]'
pytest -q -m "not network" # mocked; add `-m network` for live-API tests
planemo lint tools/manifest_gdc tools/manifest_downloader
```

## License

See [LICENSE](LICENSE).
2 changes: 1 addition & 1 deletion cli_tools/mcdi/mcdi/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@

import os

__version__ = "0.6.0"
__version__ = "0.7.0"

# Container image build id (e.g. git SHA); empty for local/editable installs.
BUILD = os.environ.get("MCDI_BUILD", "").strip()
Expand Down
24 changes: 23 additions & 1 deletion cli_tools/mcdi/mcdi/download/cli.py
Original file line number Diff line number Diff line change
Expand Up @@ -56,6 +56,13 @@ def add_arguments(subparsers: argparse._SubParsersAction) -> argparse.ArgumentPa
"--retry-backoff", type=float, default=5.0,
help="Seconds to wait before each retry pass, multiplied by the attempt number (default: 5.0)",
)
parser.add_argument(
"--retry-manifest-out", type=Path,
help="Where to write the retry manifest on a partial failure (default: "
"'<manifest>.retry<ext>', next to --manifest). Only written if some but not all files "
"failed; use this to redirect it somewhere writable when --manifest's own directory "
"isn't (e.g. a Galaxy job's staged input).",
)
parser.add_argument("--verbose", action="store_true")
parser.set_defaults(func=run)
return parser
Expand Down Expand Up @@ -113,4 +120,19 @@ def run(args: argparse.Namespace) -> int:
f"\nDone: {len(results)} total, {len(failed)} error(s), "
f"{len(mismatches)} checksum mismatch(es), {len(extract_failed)} extraction error(s)"
)
return 1 if failed or mismatches or extract_failed else 0

retry_ids = {r.entry.file_id for r in failed + mismatches}
partial = 0 < len(retry_ids) < len(results)
if partial:
retry_manifest = args.retry_manifest_out or args.manifest.with_name(
args.manifest.stem + ".retry" + args.manifest.suffix
)
source.write_subset_manifest(args.manifest, retry_ids, retry_manifest)
print(f"Wrote retry manifest for {len(retry_ids)} failed file(s) to {retry_manifest}")

if not (failed or mismatches or extract_failed):
return 0
# Some but not all files ended up missing/wrong (or all downloaded but some failed to
# extract): outputs exist but are incomplete. Distinct from 1 (nothing usable at all) so
# callers - e.g. the Galaxy wrapper's <stdio> exit-code mapping - can tell the two apart.
return 3 if len(retry_ids) < len(results) else 1
34 changes: 34 additions & 0 deletions cli_tools/mcdi/mcdi/download/sources/base.py
Original file line number Diff line number Diff line change
Expand Up @@ -28,6 +28,30 @@ def read_lines(path: Path, limit: int = CONTENT_SNIFF_LINES) -> list[str]:
return [line for line, _ in zip(f, range(limit))]


def write_csv_subset(
original_path: Path,
output_path: Path,
delimiter: str,
key_column: str,
keep_keys: set[str],
) -> None:
"""Re-parse ``original_path`` and write the rows whose ``key_column`` value is in
``keep_keys`` to ``output_path``, preserving the original header/columns and delimiter.

Used to carve a re-parseable "retry manifest" out of a manifest a service already
produced, rather than reconstructing rows from the fields ``FileEntry`` retains
(which would risk dropping columns a service's format requires but mcdi doesn't use).
"""
with open(original_path, newline="") as fin:
reader = csv.DictReader(fin, delimiter=delimiter)
fieldnames = reader.fieldnames or []
rows = [row for row in reader if (row.get(key_column) or "").strip() in keep_keys]
with open(output_path, "w", newline="") as fout:
writer = csv.DictWriter(fout, fieldnames=fieldnames, delimiter=delimiter)
writer.writeheader()
writer.writerows(rows)


@dataclass
class FileEntry:
"""A single file to download, normalized across manifest formats."""
Expand Down Expand Up @@ -63,6 +87,16 @@ def parse_manifest(self, path: Path) -> list[FileEntry]:
def request_kwargs(self, entry: FileEntry) -> dict:
"""Extra kwargs (e.g. headers) to pass to requests.get() for this entry."""

@abstractmethod
def write_subset_manifest(
self, original_manifest: Path, failed_file_ids: set[str], output_path: Path
) -> None:
"""Write a manifest at ``output_path``, in this source's own format, containing
only the rows for ``failed_file_ids`` (``FileEntry.file_id`` values) from
``original_manifest`` — so a partial download failure can be retried by feeding
``output_path`` straight back into ``mcdi download``.
"""

def rate_limit(self) -> Optional["RateLimit"]:
"""Optional pacing/rate-limit policy applied to downloads from this source."""
return None
Expand Down
9 changes: 8 additions & 1 deletion cli_tools/mcdi/mcdi/download/sources/gdc.py
Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,7 @@

import requests

from .base import FileEntry, Source
from .base import FileEntry, Source, write_csv_subset

log = logging.getLogger("mcdi.download.gdc")

Expand Down Expand Up @@ -53,6 +53,13 @@ def request_kwargs(self, entry: FileEntry) -> dict:
headers["X-Auth-Token"] = self.token
return {"headers": headers}

def write_subset_manifest(
self, original_manifest: Path, failed_file_ids: set[str], output_path: Path
) -> None:
write_csv_subset(
original_manifest, output_path, delimiter="\t", key_column="id", keep_keys=failed_file_ids
)

def known_open(self, entries: list[FileEntry], session: requests.Session) -> set[str]:
"""Bulk-query GDC's ``access`` field for every id (no token needed)."""
ids = [e.file_id for e in entries]
Expand Down
42 changes: 40 additions & 2 deletions cli_tools/mcdi/mcdi/download/sources/idc.py
Original file line number Diff line number Diff line change
Expand Up @@ -13,7 +13,7 @@

from ...errors import ApiError, InputError
from ...net import build_session
from .base import CANDIDATE_DELIMITERS, FileEntry, Source, read_header
from .base import CANDIDATE_DELIMITERS, FileEntry, Source, read_header, write_csv_subset

# Columns in an IDC cohort manifest (CSV/TSV): series-level (one row per
# series) or BigQuery-export (one row per SOPInstanceUID, same columns).
Expand Down Expand Up @@ -166,6 +166,9 @@ class IDCSource(Source):

def __init__(self, lookup: Optional[IdcIndexLookup] = None):
self._lookup = lookup or IdcIndexLookup()
# file_id (S3 key) -> (series_uid, crdc_series_uuid), populated by parse_manifest;
# lets write_subset_manifest map a failed download back to its manifest row.
self._series_by_file_id: dict[str, tuple[str, str]] = {}

@staticmethod
def sniff(header_fields: list[str]) -> bool:
Expand All @@ -176,6 +179,7 @@ def sniff_lines(lines: list[str]) -> bool:
return any(_S5CMD_CP_RE.match(line.strip()) for line in lines)

def parse_manifest(self, path: Path) -> list[FileEntry]:
self._series_by_file_id = {}
column, ids = _extract_series_refs(path)
if not ids:
raise InputError(f"IDC manifest {path} contained no resolvable series references.")
Expand Down Expand Up @@ -205,7 +209,7 @@ def _entries_for_series(self, session: requests.Session, info: SeriesInfo) -> li
Path("idc") / info.collection_id / info.patient_id / info.study_uid
/ f"{info.modality}_{info.series_uid}"
)
return [
entries = [
FileEntry(
file_id=key,
filename=key.rsplit("/", 1)[-1],
Expand All @@ -216,6 +220,9 @@ def _entries_for_series(self, session: requests.Session, info: SeriesInfo) -> li
)
for key, size in _list_series_objects(session, info.aws_bucket, info.crdc_series_uuid)
]
for entry in entries:
self._series_by_file_id[entry.file_id] = (info.series_uid, info.crdc_series_uuid)
return entries

def request_kwargs(self, entry: FileEntry) -> dict:
return {}
Expand All @@ -224,3 +231,34 @@ def known_open(self, entries: list[FileEntry], session: requests.Session) -> set
"""Already confirmed to exist via the listing in ``parse_manifest`` - skip the
redundant per-file probe in ``engine.check_access``."""
return {e.file_id for e in entries}

def write_subset_manifest(
self, original_manifest: Path, failed_file_ids: set[str], output_path: Path
) -> None:
"""Map each failed file back to the series (manifest row) it came from - one
manifest row can expand into many files, so a failure pulls in the whole series.
Re-running the result is safe: files already on disk are skipped (see
``engine._already_present``)."""
delimiter = _csv_delimiter(original_manifest)
if delimiter is not None:
series_uids = {
self._series_by_file_id[fid][0]
for fid in failed_file_ids
if fid in self._series_by_file_id
}
write_csv_subset(
original_manifest, output_path, delimiter=delimiter, key_column="SeriesInstanceUID",
keep_keys=series_uids,
)
return

crdc_uuids = {
self._series_by_file_id[fid][1]
for fid in failed_file_ids
if fid in self._series_by_file_id
}
with open(original_manifest) as fin, open(output_path, "w") as fout:
for raw in fin:
m = _S5CMD_CP_RE.match(raw.strip())
if m and m.group(1) in crdc_uuids:
fout.write(raw)
Loading
Loading