Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
27 changes: 26 additions & 1 deletion CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -69,6 +69,32 @@ and this project adheres to [Semantic Versioning](https://semver.org/).
`accumulate_arrays_are_independent_of_thread_count` checks every output
array and a `slice_by` re-accumulation bitwise across 1 to 16 threads on a
dataset with tied scores across images.
- **`COCO::create_index` reserves its (image, category) index for the number of
distinct pairs it will hold, not the number of annotations, and derives the
category-to-images index from those pairs instead of pushing once per
annotation.** `img_cat_to_anns` holds one entry per distinct `(img, cat)`
pair — on a 1.5M-annotation, 300-per-image RF-DETR-shaped workload that is
~400K pairs, so reserving for the annotation count left most of the table's
capacity unused. `cat_to_imgs` is now built from those already-unique pair
keys after the annotation loop (~400K pushes) instead of once per annotation
(~1.5M), then sorted into the same shape as before. All six index maps
(`anns`, `imgs`, `cats`, `img_to_anns`, `cat_to_imgs`, `img_cat_to_anns`) —
all private — also switched from the standard library's `HashMap` (SipHash)
to `rustc_hash::FxHashMap`, which is faster on the integer and integer-pair
keys these indices use throughout. `rustc-hash` was already in the dependency
graph transitively (via `numpy`); this makes it a direct dependency of
`hotcoco` (MIT/Apache-2.0). FxHash is not resistant to adversarially chosen
keys, an accepted trade-off for a local library indexing ids the caller
already chose to load — map iteration order was confirmed to never reach any
observable output before making the swap (every iteration site feeds a sort).
On the same RF-DETR-shaped workload, `gt.loadRes(ndarray)` is 34% faster and
the `COCOeval` constructor 49% faster (both call `create_index`); end-to-end
20% faster. `precision`, `recall`, `scores`, and `stats` are bit-identical to
before across ten configurations, including `maxDets` reassigned between
`evaluate()` and `accumulate()`. `coco::tests::test_cat_to_imgs_derived_from_pair_keys`
pins `cat_to_imgs`'s membership and deduplication and `get_ann_ids_for_img_cat`'s
dataset-order contract, and fails if the two index maps are conflated or the
order guarantee is dropped.

### Fixed

Expand Down Expand Up @@ -316,7 +342,6 @@ and this project adheres to [Semantic Versioning](https://semver.org/).
back to. Previously the marker existed only on `EvalReport`, which nothing that
writes a file uses, so comparability died with the process.


- **`hotcoco.metrics` and `hotcoco.primitives` — the functional layer.** Metric
functions you can call on plain arrays, with no evaluator, no dataset, and no COCO
JSON:
Expand Down
1 change: 1 addition & 0 deletions Cargo.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

1 change: 1 addition & 0 deletions crates/hotcoco/Cargo.toml
Original file line number Diff line number Diff line change
Expand Up @@ -21,6 +21,7 @@ rand = "0.9"
thiserror = "2"
memchr = "2.8.3"
simd-json = "0.15.1"
rustc-hash = "2"

[dev-dependencies]
tempfile = "3"
Expand Down
154 changes: 136 additions & 18 deletions crates/hotcoco/src/coco.rs
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,8 @@ use std::borrow::Cow;
use std::collections::{HashMap, HashSet};
use std::path::Path;

use rustc_hash::FxHashMap;

use crate::mask;
use crate::types::{Annotation, Category, Dataset, Image, Rle, Segmentation};

Expand All @@ -20,25 +22,25 @@ pub struct COCO {
/// [`load_warnings`](Self::load_warnings).
warnings: Vec<String>,
/// ann_id -> index into dataset.annotations
anns: HashMap<u64, usize>,
anns: FxHashMap<u64, usize>,
/// img_id -> index into dataset.images
imgs: HashMap<u64, usize>,
imgs: FxHashMap<u64, usize>,
/// cat_id -> index into dataset.categories
cats: HashMap<u64, usize>,
cats: FxHashMap<u64, usize>,
/// img_id -> [ann_id, ...]
img_to_anns: HashMap<u64, Vec<u64>>,
img_to_anns: FxHashMap<u64, Vec<u64>>,
/// cat_id -> [img_id, ...] (unique)
/// `pub(crate)` so `quality::stats` can read it — `COCO::stats` lives there,
/// since dataset statistics are introspection output rather than schema.
pub(crate) cat_to_imgs: HashMap<u64, Vec<u64>>,
pub(crate) cat_to_imgs: FxHashMap<u64, Vec<u64>>,
/// (img_id, cat_id) -> [ann_id, ...] in JSON array order.
///
/// Deliberately *not* sorted by id: pycocotools builds `_gts` by iterating
/// `dataset['annotations']` once, so array order is what feeds the matcher,
/// and the greedy tie-break (`>=`, later GT wins on equal IoU) makes that
/// order observable through `evalImgs`. Official COCO files are id-ordered
/// anyway; converted or merged files are where the two orders differ.
img_cat_to_anns: HashMap<(u64, u64), Vec<u64>>,
img_cat_to_anns: FxHashMap<(u64, u64), Vec<u64>>,
}

/// What kind of results a detection file holds, decided from its first
Expand Down Expand Up @@ -283,12 +285,12 @@ impl COCO {
let mut coco = COCO {
dataset,
warnings: Vec::new(),
anns: HashMap::new(),
imgs: HashMap::new(),
cats: HashMap::new(),
img_to_anns: HashMap::new(),
cat_to_imgs: HashMap::new(),
img_cat_to_anns: HashMap::new(),
anns: FxHashMap::default(),
imgs: FxHashMap::default(),
cats: FxHashMap::default(),
img_to_anns: FxHashMap::default(),
cat_to_imgs: FxHashMap::default(),
img_cat_to_anns: FxHashMap::default(),
};
coco.create_index();
coco
Expand Down Expand Up @@ -319,7 +321,14 @@ impl COCO {
self.cat_to_imgs.clear();
self.cat_to_imgs.reserve(n_cats);
self.img_cat_to_anns.clear();
self.img_cat_to_anns.reserve(n_anns);
// Bounded by distinct (img, cat) pairs, not annotation count — annotations
// routinely outnumber pairs by several times (e.g. ~1.5M anns over ~400K
// pairs on a dense detection workload), so reserving for `n_anns` leaves
// most of the table's capacity unused. This is a hint, not a cap: ids may
// reference images/categories absent from `images`/`categories`, and
// either list may be empty (bound 0) — the map still grows as needed.
self.img_cat_to_anns
.reserve(n_anns.min(n_imgs.saturating_mul(n_cats)));

// Single pass over annotations: build all annotation-derived indices at once
let mut dup_ann_ids = 0usize;
Expand All @@ -337,10 +346,13 @@ impl COCO {
.entry((ann.image_id, ann.category_id))
.or_default()
.push(ann.id);
self.cat_to_imgs
.entry(ann.category_id)
.or_default()
.push(ann.image_id);
}
// cat_to_imgs derived from the pair keys just built, not a second push per
// annotation: each (img, cat) key is already unique, so this is one push
// per distinct pair instead of one per annotation (~400K vs ~1.5M on the
// workload above).
for &(img_id, cat_id) in self.img_cat_to_anns.keys() {
self.cat_to_imgs.entry(cat_id).or_default().push(img_id);
}
if let Some(id) = first_dup {
// pycocotools parity: the id lookup keeps the last annotation with
Expand Down Expand Up @@ -368,7 +380,11 @@ impl COCO {
self.cats.insert(cat.id, i);
}

// Deduplicate cat_to_imgs (multiple annotations per image produce duplicates)
// Sort cat_to_imgs into the shape callers rely on (get_img_ids binary-searches
// it, stats reads its length). Each id is already unique — one push per
// distinct (img, cat) pair above — but iteration order over img_cat_to_anns's
// keys is unspecified, so the sort is still required for determinism; dedup
// stays as a cheap no-op safety net rather than an assumed invariant.
for ids in self.cat_to_imgs.values_mut() {
ids.sort_unstable();
ids.dedup();
Expand Down Expand Up @@ -1213,6 +1229,108 @@ mod tests {
assert_eq!(coco.cats.len(), 2);
}

/// `cat_to_imgs` must be derived from the distinct `(img, cat)` pairs, not
/// pushed once per annotation — img1/cat1 has three annotations (ids given
/// out of JSON order) and must collapse to one `cat_to_imgs` entry, while
/// `img_cat_to_anns` for that pair must keep every id, in dataset order.
#[test]
fn test_cat_to_imgs_derived_from_pair_keys() {
let dataset = Dataset {
info: None,
images: vec![
Image {
id: 1,
file_name: "img1.jpg".into(),
height: 100,
width: 100,
..Default::default()
},
Image {
id: 2,
file_name: "img2.jpg".into(),
height: 100,
width: 100,
..Default::default()
},
],
annotations: vec![
// img1/cat1, three annotations, ids given in descending order —
// dataset (JSON array) order is 30, 20, 10, not ascending.
Annotation {
id: 30,
image_id: 1,
category_id: 1,
..Default::default()
},
Annotation {
id: 20,
image_id: 1,
category_id: 1,
..Default::default()
},
Annotation {
id: 10,
image_id: 1,
category_id: 1,
..Default::default()
},
// img1/cat2, one annotation
Annotation {
id: 40,
image_id: 1,
category_id: 2,
..Default::default()
},
// img2/cat1, one annotation
Annotation {
id: 50,
image_id: 2,
category_id: 1,
..Default::default()
},
],
categories: vec![
Category {
id: 1,
name: "cat".into(),
..Default::default()
},
Category {
id: 2,
name: "dog".into(),
..Default::default()
},
],
licenses: vec![],
};
let coco = COCO::from_dataset(dataset);

// (1) cat_to_imgs: exact membership, three img1/cat1 annotations
// collapse to one entry, not three.
assert_eq!(coco.cat_to_imgs.get(&1), Some(&vec![1, 2]));
assert_eq!(coco.cat_to_imgs.get(&2), Some(&vec![1]));

// (2) img_cat_to_anns keeps every id, in dataset (JSON array) order —
// not sorted ascending, not deduplicated by anything upstream.
assert_eq!(
coco.get_ann_ids_for_img_cat(1, 1),
&[30, 20, 10],
"must preserve dataset order, the greedy tie-break's visibility contract"
);

// (3) the one externally observable consumer of cat_to_imgs's length.
let stats = coco.stats();
let cat1 = stats
.per_category
.iter()
.find(|c| c.id == 1)
.expect("category 1 present");
assert_eq!(
cat1.img_count, 2,
"cat 1 appears on img1 and img2, once each"
);
}

#[test]
fn test_get_ann_ids_by_img() {
let coco = COCO::from_dataset(make_test_dataset());
Expand Down
Loading