escapepod.Reader gives efficient, memory-mapped access to a single POD5
file. escapepod.DatasetReader presents many files as one stream.
import escapepod
reader = escapepod.Reader("experiment.pod5")Reader accepts a str or any os.PathLike (e.g. pathlib.Path). Use it as a
context manager so the underlying file is released promptly:
from pathlib import Path
with escapepod.Reader(Path("experiment.pod5")) as reader:
print(reader.read_count)Reader exposes file-level metadata as properties:
reader.read_count # number of reads
reader.read_batch_count # number of internal read batches
reader.signal_row_count # number of signal rows
reader.file_identifier # file UUID
reader.software # writer software string
reader.pod5_version # POD5 format version
reader.run_infos # list[RunInfo]
reader.has_index # whether a .p5s read index is available
len(reader) # same as reader.read_countIterating the reader yields ReadData objects — one per
read, metadata only:
for read in reader:
print(read.read_id, read.channel, read.num_samples, read.end_reason)reader.reads() returns the same reads as a list. Pass a selection of read
IDs to restrict it:
reads = reader.reads() # every read
subset = reader.reads(selection=["<uuid-1>", "<uuid-2>"])
subset = reader.reads(selection=ids, missing_ok=True) # skip IDs not presentBy default a requested ID that isn't in the file raises KeyError; pass
missing_ok=True to silently skip it.
reader.read_ids() returns just the IDs as a list[str].
read = reader.get_read("<uuid>") # one read, ValueError if absent
reads = reader.get_reads(["<uuid-1>", "<uuid-2>"]) # many reads
reads = reader.get_reads(ids, missing_ok=True) # skip absent IDsRepeated lookups build and reuse an in-memory index. Call
reader.build_index() up front to pay that cost once (it returns the number of
reads indexed and persists the index in the .p5s sidecar); reader.has_index
reports whether a sidecar exists.
Per-read annotations recorded with escpod annotate (e.g. demux barcode
assignments) live in the .p5s sidecar next to the POD5 and are exposed as
plain dicts:
reader.annotation_names() # e.g. ["barcode", "condition"]
barcodes = reader.annotation() # dict[read_id, label]; name= if several
conditions = reader.annotation("condition") # derived from the design
reader.design() # {"key_columns": …, "value_columns": …,
# "rows": [...]} or NoneNumeric columns (e.g. a demux classifier's confidence/margin scores) live alongside the string annotations, under their own accessors — a separate Arrow type, so a separate pair of methods:
reader.score_names() # e.g. ["ldx_confidence", "ldx_crf_margin"]
scores = reader.score("ldx_confidence") # dict[read_id, float]Unassigned/unscored reads are absent from the dict. A sidecar that does not
match the POD5 (stale, or copied from another file) raises instead of
returning wrong answers. The sidecar itself is a plain Arrow IPC table, so
pyarrow.ipc.open_file("reads.pod5.p5s").read_all() works too.
Signal is stored separately from metadata and is fetched on demand from the
reader — it is not an attribute of ReadData. Request it per read:
read = reader.reads()[0]
signal = reader.get_signal(read) # numpy int16, raw ADC values
signal_pa = reader.get_signal_pa(read) # numpy float32, picoamps (calibrated)get_signal returns raw ADC counts as int16. get_signal_pa applies the
read's calibration ((adc + offset) * scale) and returns float32 picoamps.
Both take an optional max_samples, decoding only that many leading samples
instead of the whole read (a read shorter than max_samples comes back
whole) — cheaper than slicing after the fact when only a prefix is needed:
prefix = reader.get_signal(read, max_samples=1000) # first 1000 ADC samplesFor many reads at once, the bulk variants decode in parallel and return
(read_id, signal) tuples; they take the same max_samples:
reads = reader.reads()
for read_id, signal in reader.get_signals(reads): # int16 ADC
...
for read_id, signal_pa in reader.get_signals_pa(reads): # float32 pA
...reader.prefetch_signal() is an optional hint that warms the signal region of
the file; reader.byte_count(read) reports the compressed on-disk size of a
read's signal.
Pull every read's metadata into a table in one call:
d = reader.to_dict() # dict[str, list] — column name -> values
df = reader.to_pandas() # pandas.DataFrame (requires pandas)
df = reader.to_polars() # polars.DataFrame (requires polars)All three accept the same selection / missing_ok arguments as reads().
Columns include read_id, channel, well, num_samples,
calibration_offset, calibration_scale, end_reason, and the rest of the
read metadata fields.
df = reader.to_pandas()
long_reads = df[df["num_samples"] > 100_000]
print(long_reads[["read_id", "channel", "num_samples"]])Each read is a ReadData with the POD5 read fields as read-only properties:
read.read_id # str (UUID)
read.read_number
read.start_sample
read.channel
read.well
read.pore_type
read.calibration_offset
read.calibration_scale
read.median_before
read.end_reason # e.g. "signal_positive", "mux_change", "unknown"
read.end_reason_forced
read.run_info_index # index into reader.run_infos
read.num_samples
read.signal_rows
# plus scaling/mux fields: tracked_scaling_*, predicted_scaling_*,
# num_reads_since_mux_change, time_since_mux_change, open_pore_level, ...If you already hold a raw ADC array, read.calibrate_signal_array(adc) converts
it to picoamps using that read's calibration:
adc = reader.get_signal(read)
pa = read.calibrate_signal_array(adc) # numpy float32DatasetReader reads a single file, a directory (scanned for *.pod5), or a
list mixing both, and presents every read across all files as one stream — the
escapepod analogue of pod5.DatasetReader.
# A directory (recurses by default)
with escapepod.DatasetReader("run_dir/") as ds:
print(ds.file_count, "files,", ds.read_count, "reads")
for read in ds:
signal = ds.get_signal(read)
# An explicit list of files and/or directories
ds = escapepod.DatasetReader(["a.pod5", "b.pod5", "more_reads/"])Control the directory scan with keyword arguments:
ds = escapepod.DatasetReader("run_dir/", recursive=False, pattern="*.pod5")DatasetReader offers the same reading surface as Reader —
reads()/read_ids(), to_dict()/to_pandas()/to_polars(),
get_signal()/get_signal_pa() and their bulk get_signals* forms,
byte_count(), iteration, and len() — plus paths, file_count,
read_count, and run_infos properties.