Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
38 changes: 37 additions & 1 deletion cli_tools/mcdi/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -116,7 +116,42 @@ The data commons is auto-detected from the manifest's header row; pass
| `--source {gdc,pdc}` | Skip auto-detection |
| `--workers N` | Concurrent downloads (default: 4). PDC always runs at 1 to respect its rate limit, regardless of this flag. |
| `--verify-checksum` | Verify each file's md5 against the manifest after download |
| `-x`, `--extract` | Extract recognized archives (`.tar.gz`, `.tgz`, `.tar.bz2`, `.tar.xz`, `.tar`, `.zip`, `.gz`, `.bz2`, `.xz`) in place after download, keeping only the extracted contents at the manifest's output path. Off by default. |
| `--token-file PATH` | File containing a GDC auth token, for controlled-access files |
| `--retries N` | Extra attempts for files that fail transiently within this run (default: 2) |
| `--retry-backoff SECONDS` | Wait before each retry pass, multiplied by the attempt number (default: 5.0) |

Some commons files are themselves archives (e.g. a `.tar.gz` bundle of slides).
`--extract` unpacks any recognized archive into the same directory it was
downloaded into, right after downloading it. On success, the archive itself
then moves to a sibling `<output-dir>.mcdi-archives/` directory (mirroring
`--output-dir`'s layout) — it isn't deleted, just relocated out of the way,
so `--output-dir` ends up holding only the extracted contents, not a
redundant copy of the packed archive next to them. A tool that recursively
collects everything under `--output-dir` (e.g. Galaxy's `discover_datasets`)
then only ever sees the actual extracted files. If extraction fails, the
archive is left where it was downloaded instead, so there's still something
to show for it. The archive's presence in `.mcdi-archives/` also doubles as
the idempotency marker, so reruns skip both re-downloading and
re-extracting.

**Pre-flight access check.** Before downloading anything, every file in the
manifest is probed with a cheap ranged request (skipping ones already
correctly present locally, and — for GDC — ones a single bulk lookup already
confirms are open-access). If even one file turns out to be inaccessible
(e.g. a controlled-access file without a valid token), the whole run aborts
before downloading *any* file, naming exactly which one(s) failed and why —
rather than downloading most of a large manifest only to fail on the last
file. This always runs; there's no flag to skip it.

**Retries.** Beyond the transport-level retries already built into every
request, a failed file (connection errors, `429`/`5xx`, or a checksum
mismatch — not permanent-looking failures like `401`/`403`/`404`) gets
`--retries` more whole-batch attempts, waiting `--retry-backoff × attempt`
seconds between passes. This matters most where nothing will manually rerun
the command for you on a failure — e.g. a Galaxy job, where a retried job
gets a fresh working directory, not the partial output of the failed attempt,
so anything not resolved within the one invocation is lost.

Controlled-access GDC files need an auth token, obtained by logging into the
GDC portal and downloading your token. Provide it either via the `GDC_TOKEN`
Expand All @@ -136,7 +171,8 @@ Output layout:

Re-running against the same manifest and output directory skips files already
downloaded (verifying checksums too, if `--verify-checksum` is set), so
interrupted runs can simply be re-run.
interrupted runs can simply be re-run — the pre-flight check skips them too,
so a rerun over mostly-complete output is cheap, not another full pass.

## Runtime environment

Expand Down
2 changes: 1 addition & 1 deletion cli_tools/mcdi/mcdi/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -15,7 +15,7 @@

import os

__version__ = "0.3.0"
__version__ = "0.4.0"

# Build identifier baked into the container image at build time (e.g. the git
# commit SHA). Lets you confirm the exact code a run used, even when the version
Expand Down
85 changes: 85 additions & 0 deletions cli_tools/mcdi/mcdi/download/archive.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,85 @@
"""Extraction of archives (tar/zip/gz/bz2/xz) downloaded from commons manifests."""

from __future__ import annotations

import bz2
import gzip
import lzma
import shutil
import tarfile
import zipfile
from pathlib import Path

_TAR_SUFFIXES = {
".tar.gz": "r:gz",
".tgz": "r:gz",
".tar.bz2": "r:bz2",
".tbz2": "r:bz2",
".tar.xz": "r:xz",
".txz": "r:xz",
".tar": "r:",
}
_SINGLE_FILE_OPENERS = {".gz": gzip.open, ".bz2": bz2.open, ".xz": lzma.open}


class ArchiveError(Exception):
"""An archive could not be safely extracted."""


def _tar_mode(name: str) -> str | None:
for suffix, mode in _TAR_SUFFIXES.items():
if name.endswith(suffix):
return mode
return None


def is_archive(path: Path) -> bool:
name = path.name.lower()
return _tar_mode(name) is not None or name.endswith(".zip") or name.endswith(tuple(_SINGLE_FILE_OPENERS))


def _check_member_path(member_name: str, dest_dir: Path) -> None:
"""Reject an archive member whose path would land outside ``dest_dir`` (zip-slip)."""
target = (dest_dir / member_name).resolve()
if target != dest_dir and dest_dir not in target.parents:
raise ArchiveError(f"archive member escapes destination: {member_name!r}")


def extract(path: Path) -> Path:
"""Extract ``path`` in place, into its own parent directory.

Returns the directory extracted into (tar/zip), or the decompressed file's
path (bare .gz/.bz2/.xz).
"""
dest_dir = path.parent.resolve()
name = path.name.lower()

tar_mode = _tar_mode(name)
if tar_mode is not None:
with tarfile.open(path, tar_mode) as tf:
members = tf.getmembers()
for member in members:
_check_member_path(member.name, dest_dir)
if member.issym() or member.islnk():
raise ArchiveError(f"refusing to extract link member: {member.name!r}")
if hasattr(tarfile, "data_filter"):
tf.extractall(dest_dir, members=members, filter="data")
else:
tf.extractall(dest_dir, members=members)
return dest_dir

if name.endswith(".zip"):
with zipfile.ZipFile(path) as zf:
for member_name in zf.namelist():
_check_member_path(member_name, dest_dir)
zf.extractall(dest_dir)
return dest_dir

for suffix, opener in _SINGLE_FILE_OPENERS.items():
if name.endswith(suffix):
target = path.with_name(path.name[: -len(suffix)])
with opener(path, "rb") as src, open(target, "wb") as dst:
shutil.copyfileobj(src, dst)
return target

raise ArchiveError(f"unsupported archive type: {path.name!r}")
43 changes: 41 additions & 2 deletions cli_tools/mcdi/mcdi/download/cli.py
Original file line number Diff line number Diff line change
Expand Up @@ -32,10 +32,30 @@ def add_arguments(subparsers: argparse._SubParsersAction) -> argparse.ArgumentPa
action="store_true",
help="Verify each downloaded file's md5 against the manifest",
)
parser.add_argument(
"-x",
"--extract",
action="store_true",
help="Extract recognized archives (.tar.gz, .tgz, .tar.bz2, .tar.xz, .tar, .zip, .gz, .bz2, .xz) "
"in place after download, keeping only the extracted contents at the manifest's output "
"path. On success the archive itself moves to a sibling '<output-dir>.mcdi-archives' "
"directory (not deleted); on failure it's left where it was downloaded. Off by default.",
)
parser.add_argument(
"--token-file",
help="Path to a file containing a GDC auth token (overrides GDC_TOKEN env var)",
)
parser.add_argument(
"--retries", type=int, default=2,
help="Extra attempts for files that fail transiently (network errors, 429/5xx, checksum "
"mismatches) within this run, on top of each request's own transport-level retries "
"(default: 2). Useful in contexts like a Galaxy job where nothing will rerun the "
"command for you on a retry.",
)
parser.add_argument(
"--retry-backoff", type=float, default=5.0,
help="Seconds to wait before each retry pass, multiplied by the attempt number (default: 5.0)",
)
parser.add_argument("--verbose", action="store_true")
parser.set_defaults(func=run)
return parser
Expand All @@ -59,19 +79,38 @@ def run(args: argparse.Namespace) -> int:
return 1

log.info("Detected source: %s (%d file(s))", source_name, len(entries))

log.info("Checking that all %d file(s) are accessible before downloading anything...", len(entries))
failures = engine.check_access(
entries, source, output_dir=args.output_dir, verify=args.verify_checksum, extract=args.extract,
workers=args.workers,
)
if failures:
log.error(
"%d of %d file(s) failed the pre-flight access check; aborting before downloading anything.",
len(failures), len(entries),
)
for failure in failures:
log.error(" %s (%s): %s", failure.entry.filename, failure.entry.file_id, failure.detail)
return 1

args.output_dir.mkdir(parents=True, exist_ok=True)
results = engine.run(
entries,
source,
args.output_dir,
workers=args.workers,
verify=args.verify_checksum,
extract=args.extract,
retries=args.retries,
retry_backoff=args.retry_backoff,
)

failed = [r for r in results if r.status == "error"]
mismatches = [r for r in results if r.status == "checksum_mismatch"]
extract_failed = [r for r in results if r.extract_error]
print(
f"\nDone: {len(results)} total, {len(failed)} error(s), "
f"{len(mismatches)} checksum mismatch(es)"
f"{len(mismatches)} checksum mismatch(es), {len(extract_failed)} extraction error(s)"
)
return 1 if failed or mismatches else 0
return 1 if failed or mismatches or extract_failed else 0
Loading
Loading