diff --git a/TODO.md b/TODO.md index 6ad1b569..9b2e4fc3 100644 --- a/TODO.md +++ b/TODO.md @@ -6,6 +6,7 @@ This file documents open work items. Each level-2 heading is a work item. Ever * A priority designated as P[N], on a line by itself. Priority ranges from 0 to 4, 0 being highest priority. * An effort level designated as E[N]. Effort ranges from 0 to 4, 4 being the most effort * At least one tag, indicated as !tag-name. +* Optional: a sort weight (as S[N]), which controls sort order within a priority/effort group. Default sort weight is 0. Negative sort weights are allowed. Higher sort weights will appear first. Tags can be arbitrary strings, but the most common tags are !feature, !maintenance, !bug, !docs, !lila, and !admin. !admin basically means "this involves a decision by the repo maintainer(s), it's not really a work item". @@ -22,6 +23,8 @@ This file is viewable at: https://dmorris.net/task-viewer/?file=https://raw.githubusercontent.com/agentmorris/MegaDetector/refs/heads/main/TODO.md +...which is generated with [Markdown Task Viewer](https://github.com/agentmorris/task-viewer). + # Title @@ -47,7 +50,7 @@ I got to a working nms() function that would support both import formats, but it Fix this, and remove the ultralytics NMS import. Before removing this item, consider whether the remaining functions that are still imported from the ultralytics/YOLO libraries are worth it, or whether we can (finally) remove those imports. This is the only significant utility function that is still imported. -P0 +P1 E1 @@ -182,9 +185,9 @@ E2 run_detector_batch only supports single-GPU (or single-/multi-CPU) inference. Add multi-GPU inference. This is P3 because in practice, when using manage_local_batch to create and run jobs, multi-GPU inference is handled naturally by breaking the task up into multiple lists of images. -P3 +P1 -E2 +E1 !feature @@ -277,18 +280,6 @@ E3 !integration -## Client-side RDE tool - -The [repeat detection elimination](https://github.com/agentmorris/MegaDetector/tree/main/megadetector/postprocessing/repeat_detection_elimination) pipeline currently requires stitching together a bunch of tools: python scripts, a 3P image viewer, the Windows explorer. It would be nice to integrate this into a proper client-side tool. This would also be a good opportunity to allow keeping just a couple of images from a repeat detection series; currently if you see one animal and 100 false positives in a detection group, you typically have to just keep the whole detection group and eat 100 false positives. - -P2 - -E3 - -!frontend -!feature - - ## Docs page improvements The [docs page](https://megadetector.readthedocs.io/en/latest/) is complete and up to date, but it could use a design review, updates to a more modern theme, and the addition of some more detailed information that is currently in the MegaDetector User's Guide. This is vague, I know, but basically "take a close look at the docs page and make it nicer". For my two cents, I like the styles used by [contextily](https://contextily.readthedocs.io/en/latest) and [pybowler](https://pybowler.io/docs/basics-usage). @@ -318,7 +309,7 @@ E2 It's often useful to run generic YOLO models on camera trap images, e.g. to complement MD with more fine-grained vehicle or background object classification. The MD Python package is a useful way to do this, if you want to, e.g., review the results in Timelapse, or combine them with MD/SpeciesNet results. This does not require any new code, just clear documentation. -P2 +P3 E1 @@ -475,28 +466,6 @@ E2 !optimization -## Support detector batch sizes other than 1 for video - -run_detector_batch supports batch inference (for GPUs); process_video does not. The requirement is only to support batching within a video, it's OK if an incomplete batch runs at the end of each video if it simplifies implementation. Make sure this is propagated to run_md_and_speciesnet. - -P0 - -E1 - -!feature - - -## RDE might remove custom fields within a detection object - -The RDE process loads detections into a pandas dataframe, then re-generates a new list of detections. There's no "official" scenario where detections might have custom properties, but I think this will result in the loss of custom properties. Assess this, then either fix it, remove this item (if this doesn't really happen), or update the effort/priority of this task. - -P0 - -E2 - -!bug - - ## More careful stride handling pytorch_detector uses a stride size of 64 for all 1280px models (which specifically means YOLOv5x6), and a stride size of 32 for all other models. This is true for all MD models that exist right now, but if we, for example, train YOLOv9 @ 1280px, or train a YOLOv5?6 model (where ? != "x") this heuristic would fail. @@ -592,13 +561,14 @@ E0 I'm treating all of the following as a single work item, because they're easier to tackle in a single session. -* run_md_and_speciesnet does not currently have the same checkpointing support that run_detector_batch has. The core functionality is there for the detection step, because it's built in to run_detector_batch, but this needs to be exposed to the CLI. Equivalent functionality needs to be added for the classification step. -* run_speciesnet_and_md does not currently incorporate sequence-/image-level classification smoothing. Add this. The core functionality already exists, it just needs to be added to run_md_and_speciesnet. -* Add other options from run_detector_batch (e.g. image_size, augment, detector options). No new functionality needs to be added, these can just be passed through to run_detector_batch. -* Add support for custom taxonomy lists. The core functionality already exists, it just needs to be added to run_md_and_speciesnet. -* GPU utilization is not where I would like it to be during the classification step, though I have not compared it to run_model. See whether GPU utilization goes up if I disable geofencing/rollup; if it does, push those back to the main process (which is currently just sitting idle) rather than the consumer process. -* Run one-time testing of run_md_and_speciesnet against run_model. -* Add permanent tests for run_md_and_speciesnet. +* run_md_and_speciesnet does not currently have the same checkpointing support that run_detector_batch has. The core functionality is there for the detection step, because it's built in to run_detector_batch, but this needs to be exposed to the CLI. Equivalent functionality needs to be added for the classification step. (P0) +* Add other options from run_detector_batch (e.g. image_size, augment, detector options). No new functionality needs to be added, these can just be passed through to run_detector_batch. (P0) +* GPU utilization is not where I would like it to be during the classification step, though I have not compared it to run_model. See whether GPU utilization goes up if I disable geofencing/rollup; if it does, push those back to the main process (which is currently just sitting idle) rather than the consumer process. (P0) +* Run one-time testing of run_md_and_speciesnet against run_model. (P0) +* Add permanent tests for run_md_and_speciesnet. (P0) + +* run_speciesnet_and_md does not currently incorporate sequence-/image-level classification smoothing. Add this. The core functionality already exists, it just needs to be added to run_md_and_speciesnet. (P1) +* Add support for custom taxonomy lists. The core functionality already exists, it just needs to be added to run_md_and_speciesnet. (P1) Create new work items for anything from this list that doesn't get done. @@ -706,7 +676,7 @@ When the GPU version of PyTorch is installed, but inference is run on the CPU (t P3 -E1 +E2 !bug @@ -802,9 +772,9 @@ This task is two-fold: * Assess whether map_location is supported on Apple silicon in recent versions of PyTorch, so we can eliminate the special case * Assess whether there is a performance/memory consumption benefit/cost to using map_location. -I last tried switching to use_map_location on mps devices on 2025.08.18, it did not go well. Dropping this to P3. +I last tried switching to use_map_location on mps devices on 2025.08.18, it did not go well. Dropped to P3 at the time, bumping it back to P2 now that a year has passed. -P3 +P2 E3 @@ -822,17 +792,6 @@ E0 !maintenance -## Remove unnecessary null failures from RDE output - -[repeat_detections_core](https://github.com/agentmorris/MegaDetector/blob/main/megadetector/postprocessing/repeat_detection_elimination/repeat_detections_core.py) adds an unnecessary "failure" field (set to null) for all successful images. This is not a violation of the format spec, but it's silly. This happens because this script goes through a pandas dataframe after an intermediate, then converts rows back to dicts before exporting. Fix this. The easiest fix is to just remove these prior to export. - -P1 - -E0 - -!maintenance - - ## Clean up long argument lists in run_detector_batch run_detector_batch (arguably the most important module in the repo) has super-long argument lists for basically every function. Other modules in the repo handle this by moving the relevant options to a dedicated options class. Do this for run_detector_batch. Backwards compatibility is not a huge issue as long as the CLI doesn't break. diff --git a/megadetector/detection/process_video.py b/megadetector/detection/process_video.py index 8f2107ad..0137ae00 100644 --- a/megadetector/detection/process_video.py +++ b/megadetector/detection/process_video.py @@ -18,13 +18,17 @@ import sys import argparse +from copy import deepcopy + from megadetector.detection import run_detector_batch from megadetector.utils.ct_utils import args_to_object from megadetector.utils.ct_utils import dict_to_kvp_list, parse_kvp_list from megadetector.detection.video_utils import _filename_to_frame_number from megadetector.detection.video_utils import find_videos -from megadetector.detection.video_utils import run_callback_on_frames_for_folder +from megadetector.detection.video_utils import run_callback_on_frames_for_folder_batched from megadetector.detection.run_detector import load_detector +from megadetector.detection.run_detector import is_gpu_available +from megadetector.detection.run_detector import try_download_known_detector from megadetector.postprocessing.validate_batch_results import \ ValidateBatchResultsOptions, validate_batch_results @@ -43,11 +47,6 @@ class ProcessVideoOptions: def __init__(self): #: Can be a model filename (.pt or .pb) or a model name (e.g. "MDV5A") - #: - #: Use the string "no_detection" to indicate that you only want to extract frames, - #: not run a model. If you do this, you almost definitely want to set - #: keep_extracted_frames to "True", otherwise everything in this module is a no-op. - #: I.e., there's no reason to extract frames, do nothing with them, then delete them. self.model_file = 'MDV5A' #: Video (of folder of videos) to process @@ -77,6 +76,12 @@ def __init__(self): #: getting into)... if you just want to pass smaller frames to MD, use max_width self.image_size = None + #: Batch size to use for inference. Batching only helps on a GPU, so this is + #: automatically reduced to 1 for CPU inference. Frames are batched within each video; + #: batches never span videos, so the last batch for each video is typically smaller + #: than [batch_size]. + self.batch_size = 1 + #: Enable image augmentation self.augment = False @@ -122,6 +127,9 @@ def _validate_video_options(options): if n_sampling_options_configured > 1: raise ValueError('frame_sample and time_sample are mutually exclusive') + if (options.batch_size is None) or (options.batch_size < 1): + raise ValueError('Illegal batch size {}'.format(options.batch_size)) + return True @@ -158,13 +166,45 @@ def process_videos(options): if options.verbose: print('Processing videos from input source {}'.format(options.input_video_file)) - detector = load_detector(options.model_file, - force_model_download=options.force_model_download, - detector_options=options.detector_options) + # Resolve model names (e.g. "MDV5A") to filenames, so we can tell (below) whether a GPU + # is available for this model type. We deliberately don't modify options.model_file, which + # is recorded in the output file. + model_file = try_download_known_detector(options.model_file, + force_download=options.force_model_download, + verbose=options.verbose) + + detector_options = options.detector_options + if detector_options is None: + detector_options = {} + else: + detector_options = deepcopy(detector_options) + + batch_size = options.batch_size + + # Batching only helps on a GPU, so we reduce the batch size to 1 if batch inference is + # requested on a CPU. This needs to happen before we load the model; some detectors + # compile themselves for a specific batch size at load time. + if batch_size > 1: - def frame_callback(image_np,image_id): - return detector.generate_detections_one_image(image_np, - image_id, + gpu_available = is_gpu_available(model_file, context_string='process_videos') + + if not gpu_available: + print('Batch size of {} requested, but no GPU is available, using batch size 1'.format( + batch_size)) + batch_size = 1 + + # Some detectors (currently just RF-DETR) need to know the batch size at the time the + # model is loaded; this is ignored by other detectors. + if batch_size != 1: + detector_options['batch_size'] = batch_size + + detector = load_detector(model_file, + force_model_download=False, + detector_options=detector_options) + + def frame_batch_callback(images_np,image_ids): + return detector.generate_detections_one_batch(images_np, + image_ids, detection_threshold=options.json_confidence_threshold, augment=options.augment, image_size=options.image_size, @@ -183,12 +223,14 @@ def frame_callback(image_np,image_id): video_folder = os.path.dirname(options.input_video_file) video_bn = os.path.basename(options.input_video_file) - md_results = run_callback_on_frames_for_folder(input_video_folder=video_folder, - frame_callback=frame_callback, - every_n_frames=every_n_frames_param, - verbose=options.verbose, - files_to_process_relative=[video_bn], - error_on_empty_video=options.exit_on_empty_video) + md_results = run_callback_on_frames_for_folder_batched( + input_video_folder=video_folder, + frame_batch_callback=frame_batch_callback, + batch_size=batch_size, + every_n_frames=every_n_frames_param, + verbose=options.verbose, + files_to_process_relative=[video_bn], + error_on_empty_video=options.exit_on_empty_video) else: @@ -197,12 +239,14 @@ def frame_callback(image_np,image_id): video_folder = options.input_video_file - md_results = run_callback_on_frames_for_folder(input_video_folder=options.input_video_file, - frame_callback=frame_callback, - every_n_frames=every_n_frames_param, - verbose=options.verbose, - recursive=options.recursive, - error_on_empty_video=options.exit_on_empty_video) + md_results = run_callback_on_frames_for_folder_batched( + input_video_folder=options.input_video_file, + frame_batch_callback=frame_batch_callback, + batch_size=batch_size, + every_n_frames=every_n_frames_param, + verbose=options.verbose, + recursive=options.recursive, + error_on_empty_video=options.exit_on_empty_video) # ...whether we're processing a file or a folder @@ -247,6 +291,14 @@ def frame_callback(image_np,image_id): assert frame_number not in im['frames_processed'], \ 'Received the same frame twice for video {}'.format(im['file']) + # The MD output format has no way to represent the failure of an individual + # frame, so we treat failed frames as if they hadn't been sampled at all. + if ('failure' in results_one_frame) and \ + (results_one_frame['failure'] is not None): + print('Warning: frame {} of video {} failed: {}'.format( + frame_number,video_fn,results_one_frame['failure'])) + continue + im['frames_processed'].append(frame_number) for det in results_one_frame['detections']: @@ -307,6 +359,8 @@ def options_to_command(options): cmd += ' --json_confidence_threshold ' + str(options.json_confidence_threshold) if options.frame_sample is not None: cmd += ' --frame_sample ' + str(options.frame_sample) + if (options.batch_size is not None) and (options.batch_size != 1): + cmd += ' --batch_size ' + str(options.batch_size) if options.verbose: cmd += ' --verbose' if options.detector_options is not None and len(options.detector_options) > 0: @@ -388,8 +442,7 @@ def main(): # noqa 'producing a new video with detections annotated')) parser.add_argument('model_file', type=str, - help='MegaDetector model file (.pt or .pb) or model name (e.g. "MDV5A"), '\ - 'or the string "no_detection" to run just frame extraction') + help='MegaDetector model file (.pt, .pth, or .pb) or model name (e.g. "MDV5A")') parser.add_argument('input_video_file', type=str, help='video file (or folder) to process') @@ -425,6 +478,13 @@ def main(): # noqa help=('Force image resizing to a specific integer size on the long '\ 'axis (not recommended to change this)')) + parser.add_argument('--batch_size', + type=int, + default=default_options.batch_size, + help='Batch size for GPU inference (default {}). Frames are batched '\ + 'within each video. CPU inference will ignore this and use '\ + 'batch_size=1.'.format(default_options.batch_size)) + parser.add_argument('--augment', action='store_true', help='Enable image augmentation') diff --git a/megadetector/detection/rfdetr_detector.py b/megadetector/detection/rfdetr_detector.py index 9f4b0c63..e8659d06 100644 --- a/megadetector/detection/rfdetr_detector.py +++ b/megadetector/detection/rfdetr_detector.py @@ -28,6 +28,80 @@ 'float32': torch.float32 } +#: Whether to allow TF32 for RF-DETR model loading and inference by default. +#: +#: TF32 is a reduced-precision format that PyTorch uses automatically on recent NVIDIA GPUs. +#: It's around 10% faster, but it makes RF-DETR's results depend on the batch size. In +#: practice, running the same image alone vs. in a batch of 4 shifts confidence values by up +#: to ~0.05. +#: +#: Note that this is not something we can leave to the defaults: PyTorch enables TF32 for +#: convolutions by default, and importing the rfdetr package additionally enables it for +#: matrix multiplications, for the whole process. +DEFAULT_USE_TF32 = False + + +class TF32ExecutionContext: + """ + Context manager that applies PyTorch's TF32 settings for the duration of a block, then + restores the values that were in effect when the block was entered. + + TF32 settings are global to a PyTorch process, so we apply them only around RF-DETR + model loading and inference. + """ + + def __init__(self, use_tf32): + """ + Initializes TF32ExecutionContext. + + Args: + use_tf32 (bool): whether TF32 should be enabled within this context + """ + + #: Whether TF32 should be enabled within this context + self.use_tf32 = use_tf32 + + #: The value of torch.get_float32_matmul_precision() when this context was entered + self.previous_matmul_precision = None + + #: The value of torch.backends.cudnn.allow_tf32 when this context was entered + self.previous_cudnn_allow_tf32 = None + + def apply_settings(self): + """ + Applies this context's TF32 settings. Called automatically on entry; also called + directly after importing rfdetr, which changes these settings as a side effect. + """ + + # "high" allows TF32 for matmuls, "highest" forces full fp32 + torch.set_float32_matmul_precision('high' if self.use_tf32 else 'highest') + torch.backends.cudnn.allow_tf32 = self.use_tf32 + + def __enter__(self): + """ + Stores the current TF32 settings, then applies this context's settings. + """ + + self.previous_matmul_precision = torch.get_float32_matmul_precision() + self.previous_cudnn_allow_tf32 = torch.backends.cudnn.allow_tf32 + + self.apply_settings() + + return self + + def __exit__(self, exception_type, exception_value, exception_traceback): + """ + Restores the TF32 settings that were in effect when this context was entered. + """ + + torch.set_float32_matmul_precision(self.previous_matmul_precision) + torch.backends.cudnn.allow_tf32 = self.previous_cudnn_allow_tf32 + + # Don't suppress exceptions + return False + +# ...class TF32ExecutionContext + #%% Model loading @@ -36,7 +110,8 @@ def load_model(detector_file, optimize_for_inference=False, batch_size=1, compile=None, - dtype=None): + dtype=None, + use_tf32=DEFAULT_USE_TF32): """ Load an RF-DETR model from an inference-ready .pth checkpoint via rfdetr.from_checkpoint(), which reads the architecture name ("Nano", @@ -61,12 +136,15 @@ def load_model(detector_file, dtype (str, optional): floating-point dtype used for inference, either "float16" or "float32". None means "use the rfdetr default", which is currently float32. Ignored if [optimize_for_inference] is False. + use_tf32 (bool, optional): whether to allow reduced-precision TF32 computations. + Enabling TF32 is around 10% faster, but makes results depend on the batch size; + see DEFAULT_USE_TF32. Returns: dict: dictionary with keys: - 'model': the loaded RF-DETR model - - 'model_type' (str): resolved variant class name (e.g. 'RFDETRSmall') - - 'image_size' (int): resolved inference resolution + - 'model_type' (str): model type name (e.g. 'RFDETRSmall') + - 'image_size' (int): inference resolution - 'detection_categories' (dict): mapping from string category IDs to class names """ @@ -75,71 +153,84 @@ def load_model(detector_file, 'Illegal dtype {}, dtype should be one of: {}'.format( dtype,', '.join(dtype_string_to_torch_dtype.keys())) - # The rfdetr package is not installed by default with the MegaDetector package, - # so we import it here (rather than at module scope) and print a friendly warning - # if it's not available. - try: - import rfdetr - except Exception: - print('\n\n*****\nIt looks like you are trying to run an RF-DETR model with the ' - 'MegaDetector Python package. This is supported, but the rfdetr package is not ' - 'installed by default. Run "pip install rfdetr" to install it, and try again.' - '\n*****\n\n') - raise - - assert detector_file.lower().endswith('.pth'), \ - '{} does not appear to be a compatible RF-DETR checkpoint'.format(detector_file) - - # This module uses rfdetr.from_checkpoint(), which relies on a 'model_config' field - # that was not present in checkpoints produced by early RF-DETR library versions. - print('Reading checkpoint metadata from: {}'.format(detector_file)) - checkpoint = torch.load(detector_file, weights_only=False, map_location='cpu') - if 'model_config' not in checkpoint: - raise ValueError( - "Model file '{}' is in an older format that this inference ".format(detector_file) + \ - "code does not support (missing 'model_config' metadata).") - del checkpoint - - # Load the model, letting from_checkpoint() resolve the model type and resolution. - # - # A caller-supplied image_size overrides the loaded resolution. - from_checkpoint_kwargs = {} - if image_size is not None: - from_checkpoint_kwargs['resolution'] = image_size - print('Loading model from {}...'.format(detector_file)) - model = rfdetr.from_checkpoint(detector_file, **from_checkpoint_kwargs) - - model_type = type(model).__name__ - image_size = model.model_config.resolution - print('Loaded {} at resolution {}'.format(model_type, image_size)) - - if optimize_for_inference: - - optimize_kwargs = {'batch_size':batch_size} - - # Leaving [compile] or [dtype] set to None means "use the rfdetr defaults", which - # are currently True and float32, respectively. - if compile is not None: - optimize_kwargs['compile'] = compile - if dtype is not None: - optimize_kwargs['dtype'] = dtype_string_to_torch_dtype[dtype] - - print('Optimizing loaded model for inference (batch size {}, compile {}, dtype {})'.format( - batch_size,str(compile),dtype)) - model.optimize_for_inference(**optimize_kwargs) - - # optimize_for_inference is off by default because it reportedly created - # inference errors in some environments. This comment suggests that specifying - # dtype=bfloat16 allows us to have our cake and eat it too, but this hasn't - # been tested. + # Everything from the rfdetr import through model construction runs inside a + # TF32ExecutionContext, for two reasons. First, importing rfdetr enables TF32 matmuls + # for the whole process, which would otherwise silently change the numerics of any other + # model running in this process; entering the context around the import means we put that + # setting back the way we found it on the way out. Second, optimize_for_inference() may + # trace/compile the model, which bakes in whatever precision is active at that time. + with TF32ExecutionContext(use_tf32) as tf32_context: + + # The rfdetr package is not installed by default with the MegaDetector package, + # so we import it here (rather than at module scope) and print a friendly warning + # if it's not available. + try: + import rfdetr + except Exception: + print('\n\n*****\nIt looks like you are trying to run an RF-DETR model with the ' + 'MegaDetector Python package. This is supported, but the rfdetr package is not ' + 'installed by default. Run "pip install rfdetr" to install it, and try again.' + '\n*****\n\n') + raise + + # Importing rfdetr changes the TF32 settings, so re-apply ours + tf32_context.apply_settings() + + assert detector_file.lower().endswith('.pth'), \ + '{} does not appear to be a compatible RF-DETR checkpoint'.format(detector_file) + + # This module uses rfdetr.from_checkpoint(), which relies on a 'model_config' field + # that was not present in checkpoints produced by early RF-DETR library versions. + print('Reading checkpoint metadata from: {}'.format(detector_file)) + checkpoint = torch.load(detector_file, weights_only=False, map_location='cpu') + if 'model_config' not in checkpoint: + raise ValueError( + "Model file '{}' is in an older format that this inference ".format(detector_file) + \ + "code does not support (missing 'model_config' metadata).") + del checkpoint + + # Load the model, letting from_checkpoint() resolve the model type and resolution. # - # https://github.com/roboflow/rf-detr/issues/326#issuecomment-3321838797 - # model.optimize_for_inference(batch_size=batch_size,dtype=torch.bfloat16) + # A caller-supplied image_size overrides the loaded resolution. + from_checkpoint_kwargs = {} + if image_size is not None: + from_checkpoint_kwargs['resolution'] = image_size + print('Loading model from {}...'.format(detector_file)) + model = rfdetr.from_checkpoint(detector_file, **from_checkpoint_kwargs) + + model_type = type(model).__name__ + image_size = model.model_config.resolution + print('Loaded {} at resolution {}'.format(model_type, image_size)) - elif (compile is not None) or (dtype is not None): + if optimize_for_inference: - print('Warning: the "compile" and/or "dtype" options were supplied, but ' + \ - 'optimize_for_inference is False, so they will have no effect.') + optimize_kwargs = {'batch_size':batch_size} + + # Leaving [compile] or [dtype] set to None means "use the rfdetr defaults", which + # are currently True and float32, respectively. + if compile is not None: + optimize_kwargs['compile'] = compile + if dtype is not None: + optimize_kwargs['dtype'] = dtype_string_to_torch_dtype[dtype] + + print('Optimizing loaded model for inference (batch size {}, compile {}, dtype {})'.format( + batch_size,str(compile),dtype)) + model.optimize_for_inference(**optimize_kwargs) + + # optimize_for_inference is off by default because it reportedly created + # inference errors in some environments. This comment suggests that specifying + # dtype=bfloat16 allows us to have our cake and eat it too, but this hasn't + # been tested. + # + # https://github.com/roboflow/rf-detr/issues/326#issuecomment-3321838797 + # model.optimize_for_inference(batch_size=batch_size,dtype=torch.bfloat16) + + elif (compile is not None) or (dtype is not None): + + print('Warning: the "compile" and/or "dtype" options were supplied, but ' + \ + 'optimize_for_inference is False, so they will have no effect.') + + # ...with TF32ExecutionContext(...) # Get class names from model # @@ -255,6 +346,7 @@ def __init__(self, model_path, detector_options=None, verbose=False): batch_size = 1 compile = None dtype = None + use_tf32 = DEFAULT_USE_TF32 if detector_options is not None: if ('image_size' in detector_options) and \ @@ -275,6 +367,9 @@ def __init__(self, model_path, detector_options=None, verbose=False): assert dtype in dtype_string_to_torch_dtype, \ 'Illegal dtype {}, dtype should be one of: {}'.format( dtype,', '.join(dtype_string_to_torch_dtype.keys())) + if ('use_tf32' in detector_options) and \ + (detector_options['use_tf32'] is not None): + use_tf32 = parse_bool_string(detector_options['use_tf32']) # If the caller asked for inference optimization, but didn't say anything about # compilation, don't compile. Compiling (torch.jit.trace) restricts the model to a @@ -303,6 +398,9 @@ def __init__(self, model_path, detector_options=None, verbose=False): #: compilation (torch.jit.trace) ties the model to a single batch size. self.required_batch_size = None + #: Whether TF32 is allowed during inference for this model; see DEFAULT_USE_TF32 + self.use_tf32 = use_tf32 + preprocess_only = False if (detector_options is not None) and \ ('preprocess_only' in detector_options) and \ @@ -321,7 +419,9 @@ def __init__(self, model_path, detector_options=None, verbose=False): optimize_for_inference=optimize_for_inference, batch_size=batch_size, compile=compile, - dtype=dtype) + dtype=dtype, + use_tf32=use_tf32) + self.model = model_info['model'] self.model_type = model_info['model_type'] self.image_size = model_info['image_size'] @@ -516,13 +616,17 @@ def generate_detections_one_batch(self, # Run inference. model.predict() returns a single Detections object for a single # image, or a list of Detections objects for a list of images. + # + # We apply our TF32 settings for the duration of inference; leaving TF32 enabled + # makes results depend on the batch size (see DEFAULT_USE_TF32). try: - if len(images_for_inference) == 1: - detections_list = [self.model.predict(images_for_inference[0], - threshold=detection_threshold)] - else: - detections_list = self.model.predict(images_for_inference, - threshold=detection_threshold) + with TF32ExecutionContext(self.use_tf32): + if len(images_for_inference) == 1: + detections_list = [self.model.predict(images_for_inference[0], + threshold=detection_threshold)] + else: + detections_list = self.model.predict(images_for_inference, + threshold=detection_threshold) except Exception as e: # If inference fails, mark all images in the batch as failed print('Warning: RF-DETR batch inference failed for {} images: {}'.format( diff --git a/megadetector/detection/run_md_and_speciesnet.py b/megadetector/detection/run_md_and_speciesnet.py index 9cc5d9ab..0f2d0a19 100644 --- a/megadetector/detection/run_md_and_speciesnet.py +++ b/megadetector/detection/run_md_and_speciesnet.py @@ -1055,6 +1055,7 @@ def _run_detection_step(source_folder: str, video_options.frame_sample = frame_sample video_options.time_sample = time_sample video_options.recursive = True + video_options.batch_size = detector_batch_size # Process videos process_videos(video_options) diff --git a/megadetector/detection/video_utils.py b/megadetector/detection/video_utils.py index a97ababf..667c85b5 100644 --- a/megadetector/detection/video_utils.py +++ b/megadetector/detection/video_utils.py @@ -339,6 +339,8 @@ def run_callback_on_frames(input_video_file, Calls the function frame_callback(np.array,image_id) on all (or selected) frames in [input_video_file]. + To run a callback on *batches* of frames, use run_callback_on_frames_batched(). + Args: input_video_file (str): video file to process frame_callback (function): callback to run on frames, should take an np.array and a string and @@ -470,28 +472,136 @@ def run_callback_on_frames(input_video_file, # ...def run_callback_on_frames(...) -def run_callback_on_frames_for_folder(input_video_folder, - frame_callback, - every_n_frames=None, - verbose=False, - recursive=True, - files_to_process_relative=None, - error_on_empty_video=False): +def run_callback_on_frames_batched(input_video_file, + frame_batch_callback, + batch_size, + every_n_frames=None, + verbose=False, + frames_to_process=None, + allow_empty_videos=False): """ - Calls the function frame_callback(np.array,image_id) on all (or selected) frames in - all videos in [input_video_folder]. + Calls the function frame_batch_callback(list of np.array,list of image_id) on batches of + up to [batch_size] frames from [input_video_file]. This is a wrapper around + run_callback_on_frames() that accumulates frames into batches; use that function to run a + callback on one frame at a time. Args: - input_video_folder (str): video folder to process - frame_callback (function): callback to run on frames, should take an np.array and a string and - return a single value. callback should expect two arguments: (1) a numpy array with image - data, in the typical PIL image orientation/channel order, and (2) a string identifier - for the frame, typically something like "frame0006.jpg" (even though it's not a JPEG - image, this is just an identifier for the frame). + input_video_file (str): video file to process + frame_batch_callback (function): callback to run on batches of frames, should take a list + of np.arrays and a list of strings, and should return a list of values with the same + length as the input lists. The two arguments are (1) a list of numpy arrays with image + data, in the typical PIL image orientation/channel order, and (2) a list of string + identifiers for those frames, typically something like "frame0006.jpg" (even though + they're not JPEG images, these are just identifiers for the frames). + batch_size (int): the maximum number of frames to pass to [frame_batch_callback] at a + time. The last batch for a video is typically smaller than [batch_size]. every_n_frames (int or float, optional): sample every Nth frame starting from the first frame; if this is None or 1, every frame is processed. If this is a negative value, it's - interpreted as a sampling rate in seconds, which is rounded to the nearest frame - sampling rate. + interpreted as a sampling rate in seconds, which is rounded to the nearest frame sampling + rate. Mutually exclusive with frames_to_process. + verbose (bool, optional): enable additional debug console output + frames_to_process (list of int, optional): process this specific set of frames; + mutually exclusive with every_n_frames. If all values are beyond the length + of the video, no frames are extracted. Can also be a single int, specifying + a single frame number. + allow_empty_videos (bool, optional): Just print a warning if a video appears to have no + frames (by default, this raises an Exception). + + Returns: + dict: dict with keys 'frame_filenames' (list), 'frame_rate' (float), 'results' (list), + in the same format returned by run_callback_on_frames(). 'results' contains the values + returned by the callback, flattened back out to one element per frame. + """ + + if batch_size is None: + batch_size = 1 + + if batch_size < 1: + raise ValueError('Illegal batch size {}'.format(batch_size)) + + # Frames that have been read, but not yet passed to [frame_batch_callback] + pending_images = [] + pending_frame_filenames = [] + + # Results for all the batches we've processed so far, flattened to one element per frame + batch_results = [] + + def _process_pending_batch(): + """ + Run [frame_batch_callback] on the frames in [pending_images], append the results to + [batch_results], and clear the pending lists. No-op if there are no pending frames. + """ + + if len(pending_images) == 0: + return + + results_this_batch = frame_batch_callback(pending_images,pending_frame_filenames) + + assert len(results_this_batch) == len(pending_images), \ + 'Batch callback returned {} results for {} frames in video {}'.format( + len(results_this_batch),len(pending_images),input_video_file) + + batch_results.extend(results_this_batch) + + pending_images.clear() + pending_frame_filenames.clear() + + # ...def _process_pending_batch() + + def frame_callback(image_np,image_id): + """ + Accumulate frames until we have a full batch. The value returned here is discarded; + the caller's results come from [frame_batch_callback]. + """ + + pending_images.append(image_np) + pending_frame_filenames.append(image_id) + + if len(pending_images) >= batch_size: + _process_pending_batch() + + return None + + # ...def frame_callback(...) + + to_return = run_callback_on_frames(input_video_file=input_video_file, + frame_callback=frame_callback, + every_n_frames=every_n_frames, + verbose=verbose, + frames_to_process=frames_to_process, + allow_empty_videos=allow_empty_videos) + + # Process any frames left over at the end of this video + _process_pending_batch() + + assert len(batch_results) == len(to_return['frame_filenames']), \ + 'Generated {} results for {} frames in video {}'.format( + len(batch_results),len(to_return['frame_filenames']),input_video_file) + + # Replace the per-frame results (which are all None) with the batched results + to_return['results'] = batch_results + + return to_return + +# ...def run_callback_on_frames_batched(...) + + +def _run_callback_on_frames_for_folder_core(input_video_folder, + video_callback, + verbose=False, + recursive=True, + files_to_process_relative=None, + error_on_empty_video=False): + """ + Shared implementation for run_callback_on_frames_for_folder() and + run_callback_on_frames_for_folder_batched(); those functions differ only in how they + process each video, which is what [video_callback] encapsulates. + + Args: + input_video_folder (str): video folder to process + video_callback (function): function that processes a single video, taking an absolute + video filename and returning a dict in the format returned by + run_callback_on_frames() verbose (bool, optional): enable additional debug console output recursive (bool, optional): recurse into [input_video_folder] files_to_process_relative (list, optional): only process specific relative paths @@ -500,12 +610,7 @@ def run_callback_on_frames_for_folder(input_video_folder, Returns: dict: dict with keys 'video_filenames' (list of str), 'frame_rates' (list of floats), - 'results' (list of list of dicts). 'video_filenames' will contain *relative* filenames. - 'results' is a list (one element per video) of lists (one element per frame) of whatever the - callback returns, typically (but not necessarily) dicts in the MD results format. - - For failed videos, the frame rate will be represented by -1, and "results" - will be a dict with at least the key "failure". + 'results' (list of list of dicts); see run_callback_on_frames_for_folder() for details. """ to_return = {'video_filenames':[],'frame_rates':[],'results':[]} @@ -546,12 +651,7 @@ def run_callback_on_frames_for_folder(input_video_folder, # per-image format) # # frame_filenames (list of frame IDs, i.e. synthetic filenames) - video_results = run_callback_on_frames(input_video_file=video_fn_abs, - frame_callback=frame_callback, - every_n_frames=every_n_frames, - verbose=verbose, - frames_to_process=None, - allow_empty_videos=False) + video_results = video_callback(video_fn_abs) except Exception as e: @@ -584,9 +684,133 @@ def run_callback_on_frames_for_folder(input_video_folder, return to_return +# ...def _run_callback_on_frames_for_folder_core(...) + + +def run_callback_on_frames_for_folder(input_video_folder, + frame_callback, + every_n_frames=None, + verbose=False, + recursive=True, + files_to_process_relative=None, + error_on_empty_video=False): + """ + Calls the function frame_callback(np.array,image_id) on all (or selected) frames in + all videos in [input_video_folder]. + + To run a callback on *batches* of frames, use run_callback_on_frames_for_folder_batched(). + + Args: + input_video_folder (str): video folder to process + frame_callback (function): callback to run on frames, should take an np.array and a string and + return a single value. callback should expect two arguments: (1) a numpy array with image + data, in the typical PIL image orientation/channel order, and (2) a string identifier + for the frame, typically something like "frame0006.jpg" (even though it's not a JPEG + image, this is just an identifier for the frame). + every_n_frames (int or float, optional): sample every Nth frame starting from the first frame; + if this is None or 1, every frame is processed. If this is a negative value, it's + interpreted as a sampling rate in seconds, which is rounded to the nearest frame + sampling rate. + verbose (bool, optional): enable additional debug console output + recursive (bool, optional): recurse into [input_video_folder] + files_to_process_relative (list, optional): only process specific relative paths + error_on_empty_video (bool, optional): by default, videos with errors or no valid frames + are silently stored as failures; this turns them into exceptions + + Returns: + dict: dict with keys 'video_filenames' (list of str), 'frame_rates' (list of floats), + 'results' (list of list of dicts). 'video_filenames' will contain *relative* filenames. + 'results' is a list (one element per video) of lists (one element per frame) of whatever the + callback returns, typically (but not necessarily) dicts in the MD results format. + + For failed videos, the frame rate will be represented by -1, and "results" + will be a dict with at least the key "failure". + """ + + def video_callback(video_fn_abs): + return run_callback_on_frames(input_video_file=video_fn_abs, + frame_callback=frame_callback, + every_n_frames=every_n_frames, + verbose=verbose, + frames_to_process=None, + allow_empty_videos=False) + + return _run_callback_on_frames_for_folder_core( + input_video_folder=input_video_folder, + video_callback=video_callback, + verbose=verbose, + recursive=recursive, + files_to_process_relative=files_to_process_relative, + error_on_empty_video=error_on_empty_video) + # ...def run_callback_on_frames_for_folder(...) +def run_callback_on_frames_for_folder_batched(input_video_folder, + frame_batch_callback, + batch_size, + every_n_frames=None, + verbose=False, + recursive=True, + files_to_process_relative=None, + error_on_empty_video=False): + """ + Calls the function frame_batch_callback(list of np.array,list of image_id) on batches of up + to [batch_size] frames from all videos in [input_video_folder]. Batches never span videos, + so the last batch for each video is typically smaller than [batch_size]. Use + run_callback_on_frames_for_folder() to run a callback on one frame at a time. + + Args: + input_video_folder (str): video folder to process + frame_batch_callback (function): callback to run on batches of frames, should take a list + of np.arrays and a list of strings, and should return a list of values with the same + length as the input lists. The two arguments are (1) a list of numpy arrays with image + data, in the typical PIL image orientation/channel order, and (2) a list of string + identifiers for those frames, typically something like "frame0006.jpg" (even though + they're not JPEG images, these are just identifiers for the frames). + batch_size (int): the maximum number of frames to pass to [frame_batch_callback] at a time + every_n_frames (int or float, optional): sample every Nth frame starting from the first frame; + if this is None or 1, every frame is processed. If this is a negative value, it's + interpreted as a sampling rate in seconds, which is rounded to the nearest frame + sampling rate. + verbose (bool, optional): enable additional debug console output + recursive (bool, optional): recurse into [input_video_folder] + files_to_process_relative (list, optional): only process specific relative paths + error_on_empty_video (bool, optional): by default, videos with errors or no valid frames + are silently stored as failures; this turns them into exceptions + + Returns: + dict: dict with keys 'video_filenames' (list of str), 'frame_rates' (list of floats), + 'results' (list of list of dicts), in the same format returned by + run_callback_on_frames_for_folder(). + """ + + # Validate the batch size here, rather than relying on the equivalent check in + # run_callback_on_frames_batched(); errors raised there would be caught by the per-video + # error handling below, and reported as (many) video failures. + if (batch_size is not None) and (batch_size < 1): + raise ValueError('Illegal batch size {}'.format(batch_size)) + + def video_callback(video_fn_abs): + return run_callback_on_frames_batched(input_video_file=video_fn_abs, + frame_batch_callback=frame_batch_callback, + batch_size=batch_size, + every_n_frames=every_n_frames, + verbose=verbose, + frames_to_process=None, + allow_empty_videos=False) + + return _run_callback_on_frames_for_folder_core( + input_video_folder=input_video_folder, + video_callback=video_callback, + verbose=verbose, + recursive=recursive, + files_to_process_relative=files_to_process_relative, + error_on_empty_video=error_on_empty_video) + +# ...def run_callback_on_frames_for_folder_batched(...) + + def video_to_frames(input_video_file, output_folder, overwrite=True, diff --git a/megadetector/postprocessing/load_api_results.py b/megadetector/postprocessing/load_api_results.py index 39a568fe..91d0aa34 100644 --- a/megadetector/postprocessing/load_api_results.py +++ b/megadetector/postprocessing/load_api_results.py @@ -15,25 +15,67 @@ #%% Imports -import json import os - -from typing import Optional -from collections.abc import Mapping +import json +import math +import shutil import pandas as pd from megadetector.utils.ct_utils import get_max_conf +from megadetector.utils.ct_utils import make_test_folder from megadetector.utils.ct_utils import write_json from megadetector.utils.wi_taxonomy_utils import load_md_or_speciesnet_file +#%% Constants + +#: Value used in the dataframes returned by load_api_results() to indicate that a field +#: was absent for a particular image in the source file. +#: +#: Fields that are not part of the MegaDetector output format - and even a few that are, +#: e.g. 'failure' - may be present for only some of the images in a results file. Pandas +#: fills the corresponding cells with NaN when a .json file is read into a dataframe, which +#: is indistinguishable from an explicit null in the .json file, and which forces columns +#: that contain only integers into floating-point representation. Instead, we fil those +#: cells with this sentinel value, which allows write_api_results() to omit those fields +#: and preserves the types of the values that *are* present. +#: +#: Use is_missing_field_value() rather than comparing to this value directly. +MISSING_FIELD_VALUE = '##megadetector-missing-field-value##' + + #%% Functions for loading .json results into a Pandas DataFrame, and writing back to .json -def load_api_results(api_output_path: str, normalize_paths: bool = True, - filename_replacements: Optional[Mapping[str, str]] = None, - force_forward_slashes: bool = True - ) -> tuple[pd.DataFrame, dict]: +def is_missing_field_value(v): + """ + Determines whether [v] - typically a value read from a dataframe returned by + load_api_results() - represents a field that had no value for a particular image. + + This is True for MISSING_FIELD_VALUE (used by load_api_results() for fields that were + absent for a particular image), for None (used for fields that were explicitly null in + the source file), and for NaN (which is what Pandas uses for absent fields in dataframes + that were not loaded by load_api_results()). + + Args: + v (object): the value to test + + Returns: + bool: whether [v] represents a missing field value + """ + + return (v is None) or \ + (isinstance(v,str) and (v == MISSING_FIELD_VALUE)) or \ + (isinstance(v,float) and math.isnan(v)) + +# ...def is_missing_field_value(...) + + +def load_api_results(api_output_path, + normalize_paths=True, + filename_replacements=None, + force_forward_slashes=True + ): r""" Loads json-formatted MegaDetector results to a Pandas DataFrame. @@ -47,7 +89,10 @@ def load_api_results(api_output_path: str, normalize_paths: bool = True, slashes in filenames Returns: - detection_results: pd.DataFrame, contains at least the columns ['file', 'detections','failure'] + detection_results: pd.DataFrame, contains at least the columns ['file', 'detections','failure']. + Cells corresponding to fields that were absent for a particular image in [api_output_path] + are populated with MISSING_FIELD_VALUE, rather than the NaN that Pandas would use by + default; see is_missing_field_value(). other_fields: a dict containing fields in the results other than 'images' """ @@ -88,6 +133,19 @@ def load_api_results(api_output_path: str, normalize_paths: bool = True, if 'max_detection_conf' not in im: im['max_detection_conf'] = get_max_conf(im) + # Populate fields that are absent for some images with a sentinel value, so we can tell + # them apart from fields that are explicitly null, and so that columns containing only + # integers don't get converted to floating-point. See MISSING_FIELD_VALUE. + image_field_names = {} + for im in detection_results['images']: + for field_name in im.keys(): + image_field_names[field_name] = True + + for im in detection_results['images']: + for field_name in image_field_names.keys(): + if field_name not in im: + im[field_name] = MISSING_FIELD_VALUE + # Pack the json output into a Pandas DataFrame detection_results = pd.DataFrame(detection_results['images']) @@ -96,6 +154,8 @@ def load_api_results(api_output_path: str, normalize_paths: bool = True, return detection_results, other_fields +# ...def load_api_results(...) + def write_api_results(detection_results_table, other_fields, out_path): """ @@ -114,9 +174,25 @@ def write_api_results(detection_results_table, other_fields, out_path): images = detection_results_table.to_json(orient='records', double_precision=3) images = json.loads(images) + for im in images: - if 'failure' in im and im['failure'] is None: + + # Remove fields that weren't present for this image in the file this table was + # loaded from; see MISSING_FIELD_VALUE. Fields that were explicitly null are + # left alone. + field_names_to_remove = [] + for field_name in im.keys(): + if isinstance(im[field_name],str) and (im[field_name] == MISSING_FIELD_VALUE): + field_names_to_remove.append(field_name) + for field_name in field_names_to_remove: + del im[field_name] + + # An explicitly-null failure indicator is meaningless, remove it + if ('failure' in im) and (im['failure'] is None): del im['failure'] + + # ...for each image + fields['images'] = images # Convert the 'version' field back to a string as per format convention @@ -144,6 +220,8 @@ def write_api_results(detection_results_table, other_fields, out_path): print('Finished writing detection results to {}'.format(out_path)) +# ...def write_api_results(...) + def load_api_results_csv(filename, normalize_paths=True, filename_replacements=None, nrows=None): """ @@ -198,6 +276,8 @@ def load_api_results_csv(filename, normalize_paths=True, filename_replacements=N return detection_results +# ...def load_api_results_csv(...) + def write_api_results_csv(detection_results, filename): """ @@ -221,3 +301,88 @@ def write_api_results_csv(detection_results, filename): detection_results.to_csv(filename, index=False) print('Finished writing detection results to {}'.format(filename)) + +# ...def write_api_results_csv(...) + + +#%% Tests + +def test_load_api_results(): + """ + Test that a .json results file survives a load_api_results()/write_api_results() + round trip, particularly fields that are present for only some images. + """ + + test_folder = make_test_folder(subfolder='load_api_results_tests') + + try: + + input_file = os.path.join(test_folder,'test_results.json') + output_file = os.path.join(test_folder,'test_results_filtered.json') + + images = [ + { + 'file':'test_folder/failure.jpg', + 'detections':None, + 'failure':'synthetic failure' + }, + { + 'file':'test_folder/string_field.jpg', + 'detections':[{'category':'1','conf':0.797, + 'bbox':[0.591,0.077,0.047,0.047]}], + 'synthetic_string_field':'synthetic value' + }, + { + 'file':'test_folder/int_field.jpg', + 'detections':[{'category':'1','conf':0.254, + 'bbox':[0.622,0.077,0.012,0.032]}], + 'synthetic_int_field':10 + }, + { + 'file':'test_folder/null_field.jpg', + 'detections':[], + 'synthetic_string_field':None + }, + { + 'file':'test_folder/no_extra_fields.jpg', + 'detections':[] + } + ] + + input_data = { + 'info':{'format_version':'1.3','detector':'synthetic_detector'}, + 'detection_categories':{'1':'animal'}, + 'synthetic_file_level_field':{'test_key':'test value'}, + 'images':images + } + + write_json(input_file,input_data) + + detection_results_table, other_fields = load_api_results(input_file) + + # Fields that are absent for an image should be represented with a sentinel value, + # rather than with the NaN Pandas would use by default + absent_field_value = detection_results_table['synthetic_string_field'].iloc[4] + assert absent_field_value == MISSING_FIELD_VALUE, \ + 'Absent field represented as {}, expected a sentinel value'.format(absent_field_value) + assert is_missing_field_value(absent_field_value), \ + 'Sentinel value not recognized as a missing field value' + + write_api_results(detection_results_table,other_fields,output_file) + + with open(output_file,'r') as f: + output_data = json.load(f) + + # Fields that were absent should still be absent, fields that were explicitly null + # should still be null, and integers should not have become floating-point values + assert output_data['images'] == images, \ + 'Image fields did not survive a load/write round trip' + assert output_data['synthetic_file_level_field'] == \ + input_data['synthetic_file_level_field'], \ + 'File-level fields did not survive a load/write round trip' + + finally: + + shutil.rmtree(test_folder,ignore_errors=True) + +# ...def test_load_api_results(...) diff --git a/megadetector/postprocessing/postprocess_batch_results.py b/megadetector/postprocessing/postprocess_batch_results.py index 4f7c6d57..121e08ba 100644 --- a/megadetector/postprocessing/postprocess_batch_results.py +++ b/megadetector/postprocessing/postprocess_batch_results.py @@ -59,6 +59,7 @@ from megadetector.data_management.cct_json_utils import CameraTrapJsonUtils from megadetector.data_management.cct_json_utils import IndexedJsonDb from megadetector.postprocessing.load_api_results import load_api_results +from megadetector.postprocessing.load_api_results import is_missing_field_value from megadetector.detection.run_detector import get_typical_confidence_threshold_from_results warnings.filterwarnings('ignore', '(Possibly )?corrupt EXIF data', UserWarning) @@ -859,9 +860,8 @@ def _render_image_no_gt(file_info, field_value = file_info[field_name] - if (field_value is None) or \ - (isinstance(field_value,float) and np.isnan(field_value)): - continue + if is_missing_field_value(field_value): + continue # Optionally use a display name that's different from the field name if isinstance(options.additional_image_fields_to_display,dict): @@ -1150,11 +1150,14 @@ def process_batch_results(options): # Remove rows with inference failures (typically due to corrupt images) n_failures = 0 if 'failure' in detections_df.columns: - n_failures = detections_df['failure'].count() + # Images without a failure indicator have a sentinel value in this column, rather + # than NaN; see load_api_results(). + b_image_succeeded = detections_df['failure'].apply(is_missing_field_value) + n_failures = int((~b_image_succeeded).sum()) print('Ignoring {} failed images'.format(n_failures)) # Explicitly forcing a copy() operation here to suppress "trying to be set # on a copy" warnings (and associated risks) below. - detections_df = detections_df[detections_df['failure'].isna()].copy() + detections_df = detections_df[b_image_succeeded].copy() assert other_fields is not None diff --git a/megadetector/postprocessing/repeat_detection_elimination/README.md b/megadetector/postprocessing/repeat_detection_elimination/README.md index 59ab92b3..c001fe95 100644 --- a/megadetector/postprocessing/repeat_detection_elimination/README.md +++ b/megadetector/postprocessing/repeat_detection_elimination/README.md @@ -18,6 +18,8 @@ This document shows you how to run these scripts. None of this is required; you can work with MegaDetector results without doing this step. In fact, we only usually recommend this if you have (a) lots of images (millions) and (b) a reasonably high rate of false positives. But if you have (a) and (b), this process may save you lots of time! +If you prefer an app to the approach presented on this page, you might want to check out Yunxuan Chai's "[RocksBeGone](https://rocksbegone.camtra.pw/)", a graphical version of the repeat detection elimination process that runs in a browser. + # Prerequisites diff --git a/megadetector/postprocessing/repeat_detection_elimination/repeat_detections_core.py b/megadetector/postprocessing/repeat_detection_elimination/repeat_detections_core.py index 8c9cb9ea..20892d03 100644 --- a/megadetector/postprocessing/repeat_detection_elimination/repeat_detections_core.py +++ b/megadetector/postprocessing/repeat_detection_elimination/repeat_detections_core.py @@ -38,6 +38,7 @@ from megadetector.utils import path_utils from megadetector.utils import ct_utils from megadetector.postprocessing.load_api_results import load_api_results, write_api_results +from megadetector.postprocessing.load_api_results import is_missing_field_value from megadetector.postprocessing.postprocess_batch_results import is_sas_url from megadetector.postprocessing.postprocess_batch_results import relative_sas_url from megadetector.visualization.visualization_utils import open_image, render_detection_bounding_boxes @@ -602,8 +603,10 @@ def _find_matches_in_directory(dir_name_and_rows, options): print('Loading results for location {} from {}'.format( dir_name,detections_loaded_from_csv_file)) rows = pd.read_csv(detections_loaded_from_csv_file) - # Pandas writes out detections out as strings, convert them back to lists - rows['detections'] = rows['detections'].apply(lambda s: json.loads(s.replace('\'','"'))) + # Pandas writes out detections out as strings, convert them back to lists. Images + # for which no detections are available (e.g. failed images) are left alone. + rows['detections'] = rows['detections'].apply( + lambda s: s if is_missing_field_value(s) else json.loads(s.replace('\'','"'))) if options.maxImagesPerFolder is not None and len(rows) > options.maxImagesPerFolder: print('Ignoring directory {} because it has {} images (limit set to {})'.format( @@ -658,8 +661,8 @@ def _find_matches_in_directory(dir_name_and_rows, options): # # } detections = row['detections'] - if isinstance(detections,float): - assert isinstance(row['failure'],str), 'Expected failure indicator' + if is_missing_field_value(detections): + assert not is_missing_field_value(row['failure']), 'Expected failure indicator' print('Skipping failed image {} ({})'.format(filename,row['failure'])) continue @@ -900,9 +903,9 @@ def _update_detection_table(repeat_detection_results, options, output_file_name= for i_row, row in detection_results.iterrows(): detections = row['detections'] - if (detections is None) or isinstance(detections,float): - assert isinstance(row['failure'],str), \ - 'Illegal failure indicator of type {}'.format(type(row['failure'])) + if is_missing_field_value(detections): + assert not is_missing_field_value(row['failure']), \ + 'Illegal failure indicator {}'.format(row['failure']) continue if len(detections) == 0: diff --git a/megadetector/utils/md_tests.py b/megadetector/utils/md_tests.py index b5d1ec30..ddd9c0ea 100644 --- a/megadetector/utils/md_tests.py +++ b/megadetector/utils/md_tests.py @@ -1078,6 +1078,25 @@ def run_python_tests(options): compare_results(video_options.output_json_file,expected_results_file,options_loose) + + ## Video test (folder, batch size > 1) + + # Batching is disabled on the CPU, so on a CPU-only machine, this is just a + # repeat of the previous test. + print('\n** Running MD on a folder of videos with batch size > 1 (module) **\n') + + video_options_batch = deepcopy(video_options) + video_options_batch.output_json_file = \ + insert_before_extension(video_options.output_json_file,'batch') + video_options_batch.batch_size = options.alternative_batch_size + + _ = process_videos(video_options_batch) + + assert os.path.isfile(video_options_batch.output_json_file), \ + 'Batched video test failed to render output .json file' + + compare_results(video_options_batch.output_json_file,expected_results_file,options_loose) + # ...if we're not skipping video tests print('\n*** Finished module tests ***\n')