Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
77 changes: 18 additions & 59 deletions TODO.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,7 @@ This file documents open work items. Each level-2 heading is a work item. Ever
* A priority designated as P[N], on a line by itself. Priority ranges from 0 to 4, 0 being highest priority.
* An effort level designated as E[N]. Effort ranges from 0 to 4, 4 being the most effort
* At least one tag, indicated as !tag-name.
* Optional: a sort weight (as S[N]), which controls sort order within a priority/effort group. Default sort weight is 0. Negative sort weights are allowed. Higher sort weights will appear first.

Tags can be arbitrary strings, but the most common tags are !feature, !maintenance, !bug, !docs, !lila, and !admin. !admin basically means "this involves a decision by the repo maintainer(s), it's not really a work item".

Expand All @@ -22,6 +23,8 @@ This file is viewable at:

https://dmorris.net/task-viewer/?file=https://raw.githubusercontent.com/agentmorris/MegaDetector/refs/heads/main/TODO.md

...which is generated with [Markdown Task Viewer](https://github.com/agentmorris/task-viewer).


# Title

Expand All @@ -47,7 +50,7 @@ I got to a working nms() function that would support both import formats, but it

Fix this, and remove the ultralytics NMS import. Before removing this item, consider whether the remaining functions that are still imported from the ultralytics/YOLO libraries are worth it, or whether we can (finally) remove those imports. This is the only significant utility function that is still imported.

P0
P1

E1

Expand Down Expand Up @@ -182,9 +185,9 @@ E2

run_detector_batch only supports single-GPU (or single-/multi-CPU) inference. Add multi-GPU inference. This is P3 because in practice, when using manage_local_batch to create and run jobs, multi-GPU inference is handled naturally by breaking the task up into multiple lists of images.

P3
P1

E2
E1

!feature

Expand Down Expand Up @@ -277,18 +280,6 @@ E3
!integration


## Client-side RDE tool

The [repeat detection elimination](https://github.com/agentmorris/MegaDetector/tree/main/megadetector/postprocessing/repeat_detection_elimination) pipeline currently requires stitching together a bunch of tools: python scripts, a 3P image viewer, the Windows explorer. It would be nice to integrate this into a proper client-side tool. This would also be a good opportunity to allow keeping just a couple of images from a repeat detection series; currently if you see one animal and 100 false positives in a detection group, you typically have to just keep the whole detection group and eat 100 false positives.

P2

E3

!frontend
!feature


## Docs page improvements

The [docs page](https://megadetector.readthedocs.io/en/latest/) is complete and up to date, but it could use a design review, updates to a more modern theme, and the addition of some more detailed information that is currently in the MegaDetector User's Guide. This is vague, I know, but basically "take a close look at the docs page and make it nicer". For my two cents, I like the styles used by [contextily](https://contextily.readthedocs.io/en/latest) and [pybowler](https://pybowler.io/docs/basics-usage).
Expand Down Expand Up @@ -318,7 +309,7 @@ E2

It's often useful to run generic YOLO models on camera trap images, e.g. to complement MD with more fine-grained vehicle or background object classification. The MD Python package is a useful way to do this, if you want to, e.g., review the results in Timelapse, or combine them with MD/SpeciesNet results. This does not require any new code, just clear documentation.

P2
P3

E1

Expand Down Expand Up @@ -475,28 +466,6 @@ E2
!optimization


## Support detector batch sizes other than 1 for video

run_detector_batch supports batch inference (for GPUs); process_video does not. The requirement is only to support batching within a video, it's OK if an incomplete batch runs at the end of each video if it simplifies implementation. Make sure this is propagated to run_md_and_speciesnet.

P0

E1

!feature


## RDE might remove custom fields within a detection object

The RDE process loads detections into a pandas dataframe, then re-generates a new list of detections. There's no "official" scenario where detections might have custom properties, but I think this will result in the loss of custom properties. Assess this, then either fix it, remove this item (if this doesn't really happen), or update the effort/priority of this task.

P0

E2

!bug


## More careful stride handling

pytorch_detector uses a stride size of 64 for all 1280px models (which specifically means YOLOv5x6), and a stride size of 32 for all other models. This is true for all MD models that exist right now, but if we, for example, train YOLOv9 @ 1280px, or train a YOLOv5?6 model (where ? != "x") this heuristic would fail.
Expand Down Expand Up @@ -592,13 +561,14 @@ E0

I'm treating all of the following as a single work item, because they're easier to tackle in a single session.

* run_md_and_speciesnet does not currently have the same checkpointing support that run_detector_batch has. The core functionality is there for the detection step, because it's built in to run_detector_batch, but this needs to be exposed to the CLI. Equivalent functionality needs to be added for the classification step.
* run_speciesnet_and_md does not currently incorporate sequence-/image-level classification smoothing. Add this. The core functionality already exists, it just needs to be added to run_md_and_speciesnet.
* Add other options from run_detector_batch (e.g. image_size, augment, detector options). No new functionality needs to be added, these can just be passed through to run_detector_batch.
* Add support for custom taxonomy lists. The core functionality already exists, it just needs to be added to run_md_and_speciesnet.
* GPU utilization is not where I would like it to be during the classification step, though I have not compared it to run_model. See whether GPU utilization goes up if I disable geofencing/rollup; if it does, push those back to the main process (which is currently just sitting idle) rather than the consumer process.
* Run one-time testing of run_md_and_speciesnet against run_model.
* Add permanent tests for run_md_and_speciesnet.
* run_md_and_speciesnet does not currently have the same checkpointing support that run_detector_batch has. The core functionality is there for the detection step, because it's built in to run_detector_batch, but this needs to be exposed to the CLI. Equivalent functionality needs to be added for the classification step. (P0)
* Add other options from run_detector_batch (e.g. image_size, augment, detector options). No new functionality needs to be added, these can just be passed through to run_detector_batch. (P0)
* GPU utilization is not where I would like it to be during the classification step, though I have not compared it to run_model. See whether GPU utilization goes up if I disable geofencing/rollup; if it does, push those back to the main process (which is currently just sitting idle) rather than the consumer process. (P0)
* Run one-time testing of run_md_and_speciesnet against run_model. (P0)
* Add permanent tests for run_md_and_speciesnet. (P0)

* run_speciesnet_and_md does not currently incorporate sequence-/image-level classification smoothing. Add this. The core functionality already exists, it just needs to be added to run_md_and_speciesnet. (P1)
* Add support for custom taxonomy lists. The core functionality already exists, it just needs to be added to run_md_and_speciesnet. (P1)

Create new work items for anything from this list that doesn't get done.

Expand Down Expand Up @@ -706,7 +676,7 @@ When the GPU version of PyTorch is installed, but inference is run on the CPU (t

P3

E1
E2

!bug

Expand Down Expand Up @@ -802,9 +772,9 @@ This task is two-fold:
* Assess whether map_location is supported on Apple silicon in recent versions of PyTorch, so we can eliminate the special case
* Assess whether there is a performance/memory consumption benefit/cost to using map_location.

I last tried switching to use_map_location on mps devices on 2025.08.18, it did not go well. Dropping this to P3.
I last tried switching to use_map_location on mps devices on 2025.08.18, it did not go well. Dropped to P3 at the time, bumping it back to P2 now that a year has passed.

P3
P2

E3

Expand All @@ -822,17 +792,6 @@ E0
!maintenance


## Remove unnecessary null failures from RDE output

[repeat_detections_core](https://github.com/agentmorris/MegaDetector/blob/main/megadetector/postprocessing/repeat_detection_elimination/repeat_detections_core.py) adds an unnecessary "failure" field (set to null) for all successful images. This is not a violation of the format spec, but it's silly. This happens because this script goes through a pandas dataframe after an intermediate, then converts rows back to dicts before exporting. Fix this. The easiest fix is to just remove these prior to export.

P1

E0

!maintenance


## Clean up long argument lists in run_detector_batch

run_detector_batch (arguably the most important module in the repo) has super-long argument lists for basically every function. Other modules in the repo handle this by moving the relevant options to a dedicated options class. Do this for run_detector_batch. Backwards compatibility is not a huge issue as long as the CLI doesn't break.
Expand Down
112 changes: 86 additions & 26 deletions megadetector/detection/process_video.py
Original file line number Diff line number Diff line change
Expand Up @@ -18,13 +18,17 @@
import sys
import argparse

from copy import deepcopy

from megadetector.detection import run_detector_batch
from megadetector.utils.ct_utils import args_to_object
from megadetector.utils.ct_utils import dict_to_kvp_list, parse_kvp_list
from megadetector.detection.video_utils import _filename_to_frame_number
from megadetector.detection.video_utils import find_videos
from megadetector.detection.video_utils import run_callback_on_frames_for_folder
from megadetector.detection.video_utils import run_callback_on_frames_for_folder_batched
from megadetector.detection.run_detector import load_detector
from megadetector.detection.run_detector import is_gpu_available
from megadetector.detection.run_detector import try_download_known_detector
from megadetector.postprocessing.validate_batch_results import \
ValidateBatchResultsOptions, validate_batch_results

Expand All @@ -43,11 +47,6 @@ class ProcessVideoOptions:
def __init__(self):

#: Can be a model filename (.pt or .pb) or a model name (e.g. "MDV5A")
#:
#: Use the string "no_detection" to indicate that you only want to extract frames,
#: not run a model. If you do this, you almost definitely want to set
#: keep_extracted_frames to "True", otherwise everything in this module is a no-op.
#: I.e., there's no reason to extract frames, do nothing with them, then delete them.
self.model_file = 'MDV5A'

#: Video (of folder of videos) to process
Expand Down Expand Up @@ -77,6 +76,12 @@ def __init__(self):
#: getting into)... if you just want to pass smaller frames to MD, use max_width
self.image_size = None

#: Batch size to use for inference. Batching only helps on a GPU, so this is
#: automatically reduced to 1 for CPU inference. Frames are batched within each video;
#: batches never span videos, so the last batch for each video is typically smaller
#: than [batch_size].
self.batch_size = 1

#: Enable image augmentation
self.augment = False

Expand Down Expand Up @@ -122,6 +127,9 @@ def _validate_video_options(options):
if n_sampling_options_configured > 1:
raise ValueError('frame_sample and time_sample are mutually exclusive')

if (options.batch_size is None) or (options.batch_size < 1):
raise ValueError('Illegal batch size {}'.format(options.batch_size))

return True


Expand Down Expand Up @@ -158,13 +166,45 @@ def process_videos(options):
if options.verbose:
print('Processing videos from input source {}'.format(options.input_video_file))

detector = load_detector(options.model_file,
force_model_download=options.force_model_download,
detector_options=options.detector_options)
# Resolve model names (e.g. "MDV5A") to filenames, so we can tell (below) whether a GPU
# is available for this model type. We deliberately don't modify options.model_file, which
# is recorded in the output file.
model_file = try_download_known_detector(options.model_file,
force_download=options.force_model_download,
verbose=options.verbose)

detector_options = options.detector_options
if detector_options is None:
detector_options = {}
else:
detector_options = deepcopy(detector_options)

batch_size = options.batch_size

# Batching only helps on a GPU, so we reduce the batch size to 1 if batch inference is
# requested on a CPU. This needs to happen before we load the model; some detectors
# compile themselves for a specific batch size at load time.
if batch_size > 1:

def frame_callback(image_np,image_id):
return detector.generate_detections_one_image(image_np,
image_id,
gpu_available = is_gpu_available(model_file, context_string='process_videos')

if not gpu_available:
print('Batch size of {} requested, but no GPU is available, using batch size 1'.format(
batch_size))
batch_size = 1

# Some detectors (currently just RF-DETR) need to know the batch size at the time the
# model is loaded; this is ignored by other detectors.
if batch_size != 1:
detector_options['batch_size'] = batch_size

detector = load_detector(model_file,
force_model_download=False,
detector_options=detector_options)

def frame_batch_callback(images_np,image_ids):
return detector.generate_detections_one_batch(images_np,
image_ids,
detection_threshold=options.json_confidence_threshold,
augment=options.augment,
image_size=options.image_size,
Expand All @@ -183,12 +223,14 @@ def frame_callback(image_np,image_id):

video_folder = os.path.dirname(options.input_video_file)
video_bn = os.path.basename(options.input_video_file)
md_results = run_callback_on_frames_for_folder(input_video_folder=video_folder,
frame_callback=frame_callback,
every_n_frames=every_n_frames_param,
verbose=options.verbose,
files_to_process_relative=[video_bn],
error_on_empty_video=options.exit_on_empty_video)
md_results = run_callback_on_frames_for_folder_batched(
input_video_folder=video_folder,
frame_batch_callback=frame_batch_callback,
batch_size=batch_size,
every_n_frames=every_n_frames_param,
verbose=options.verbose,
files_to_process_relative=[video_bn],
error_on_empty_video=options.exit_on_empty_video)

else:

Expand All @@ -197,12 +239,14 @@ def frame_callback(image_np,image_id):

video_folder = options.input_video_file

md_results = run_callback_on_frames_for_folder(input_video_folder=options.input_video_file,
frame_callback=frame_callback,
every_n_frames=every_n_frames_param,
verbose=options.verbose,
recursive=options.recursive,
error_on_empty_video=options.exit_on_empty_video)
md_results = run_callback_on_frames_for_folder_batched(
input_video_folder=options.input_video_file,
frame_batch_callback=frame_batch_callback,
batch_size=batch_size,
every_n_frames=every_n_frames_param,
verbose=options.verbose,
recursive=options.recursive,
error_on_empty_video=options.exit_on_empty_video)

# ...whether we're processing a file or a folder

Expand Down Expand Up @@ -247,6 +291,14 @@ def frame_callback(image_np,image_id):
assert frame_number not in im['frames_processed'], \
'Received the same frame twice for video {}'.format(im['file'])

# The MD output format has no way to represent the failure of an individual
# frame, so we treat failed frames as if they hadn't been sampled at all.
if ('failure' in results_one_frame) and \
(results_one_frame['failure'] is not None):
print('Warning: frame {} of video {} failed: {}'.format(
frame_number,video_fn,results_one_frame['failure']))
continue

im['frames_processed'].append(frame_number)

for det in results_one_frame['detections']:
Expand Down Expand Up @@ -307,6 +359,8 @@ def options_to_command(options):
cmd += ' --json_confidence_threshold ' + str(options.json_confidence_threshold)
if options.frame_sample is not None:
cmd += ' --frame_sample ' + str(options.frame_sample)
if (options.batch_size is not None) and (options.batch_size != 1):
cmd += ' --batch_size ' + str(options.batch_size)
if options.verbose:
cmd += ' --verbose'
if options.detector_options is not None and len(options.detector_options) > 0:
Expand Down Expand Up @@ -388,8 +442,7 @@ def main(): # noqa
'producing a new video with detections annotated'))

parser.add_argument('model_file', type=str,
help='MegaDetector model file (.pt or .pb) or model name (e.g. "MDV5A"), '\
'or the string "no_detection" to run just frame extraction')
help='MegaDetector model file (.pt, .pth, or .pb) or model name (e.g. "MDV5A")')

parser.add_argument('input_video_file', type=str,
help='video file (or folder) to process')
Expand Down Expand Up @@ -425,6 +478,13 @@ def main(): # noqa
help=('Force image resizing to a specific integer size on the long '\
'axis (not recommended to change this)'))

parser.add_argument('--batch_size',
type=int,
default=default_options.batch_size,
help='Batch size for GPU inference (default {}). Frames are batched '\
'within each video. CPU inference will ignore this and use '\
'batch_size=1.'.format(default_options.batch_size))

parser.add_argument('--augment',
action='store_true',
help='Enable image augmentation')
Expand Down
Loading
Loading