Skip to content

enhance: add batch validity APIs to IdMap - #1858

Merged
sre-ci-robot merged 1 commit into
zilliztech:mainfrom
marcelo-cjl:codex/batch-idmap-validity
Sep 29, 2026
Merged

sre-ci-robot merged 1 commit into
zilliztech:mainfrom
marcelo-cjl:codex/batch-idmap-validity

Conversation

@marcelo-cjl

@marcelo-cjl marcelo-cjl commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • Add snapshot-based range and arbitrary-offset validity visitors to IdMap.
  • Add a packed range API that ANDs validity directly into a caller-provided bitmap, including non-byte-aligned source and target ranges.
  • Process contiguous bitmap bytes with AVX2/SSE2 on x86 and NEON on ARM, with a scalar fallback; growing storage performs one initial record lookup and then traverses records sequentially.
  • Preserve the existing visibility snapshot semantics and cover record boundaries, negative/out-of-range IDs, disabled maps, and append-during-visit consistency.
  • Follow up on the batch-validity performance review from fix: support vector null predicates milvus-io/milvus#52442.

Performance

Kernel microbenchmark over 4,194,304 bits per invocation (median of 9 retained rounds; non-byte-aligned source/target):

  • x86 SSE2: packed path is 405x faster than the sequential bool visitor for sealed storage and 137x for growing storage.
  • x86 AVX2: 525x faster for sealed storage and 150x for growing storage.
  • ARM64 NEON (AWS m8g): 404x faster for sealed storage and 139x for growing storage.

These are cache-resident kernel measurements rather than end-to-end Milvus latency.

Test Plan

  • Release build: make WITH_UT=True BUILD_DIR=build/main/cpu-ut CONAN_INSTALL_FLAGS='--build=missing --build=liburing'
  • build/main/cpu-ut/Release/tests/ut/knowhere_tests '[nullable][id_map]' (7 test cases, 1137 assertions)
  • Standalone packed-path correctness and benchmark on x86 SSE2 and AVX2
  • Standalone packed-path correctness and benchmark on ARM64 NEON (AWS m8g)
  • pre-commit run --all-files --show-diff-on-failure

@mergify

mergify Bot commented Sep 29, 2026

Copy link
Copy Markdown

@marcelo-cjl 🔍 Important: PR Classification Needed!

For efficient project management and a seamless review process, it's essential to classify your PR correctly. Here's how:

  1. If you're fixing a bug, label it as kind/bug.
  2. For small tweaks (less than 20 lines without altering any functionality), please use kind/improvement.
  3. Significant changes that don't modify existing functionalities should be tagged as kind/enhancement.
  4. Adjusting APIs or changing functionality? Go with kind/feature.

For any PR outside the kind/improvement category, ensure you link to the associated issue using the format: “issue: #”.

Thanks for your efforts and contribution to the community!.

@marcelo-cjl marcelo-cjl changed the title feat: add batch validity APIs to IdMap enhance: add batch validity APIs to IdMap Sep 29, 2026
@marcelo-cjl marcelo-cjl changed the title enhance: add batch validity APIs to IdMap [WIP]enhance: add batch validity APIs to IdMap Sep 29, 2026
@marcelo-cjl
marcelo-cjl force-pushed the codex/batch-idmap-validity branch from 40e4827 to a665560 Compare September 29, 2026 03:47
@marcelo-cjl marcelo-cjl changed the title [WIP]enhance: add batch validity APIs to IdMap feat: add batch validity APIs to IdMap Sep 29, 2026
Signed-off-by: marcelo-cjl <marcelo.chen@zilliz.com>
@marcelo-cjl
marcelo-cjl force-pushed the codex/batch-idmap-validity branch from a665560 to ecb0b4c Compare September 29, 2026 03:50
@marcelo-cjl marcelo-cjl changed the title feat: add batch validity APIs to IdMap enhance: add batch validity APIs to IdMap Sep 29, 2026
@mergify mergify Bot added the ci-passed label Sep 29, 2026
// non-zero. x86 and AArch64 use their baseline vector ISA, with AVX2 selected
// when the including translation unit is compiled for it.
inline void
AndPackedBytes(uint8_t* target, const uint8_t* source, size_t byte_count, unsigned shift) {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I would prefer to have distinct versions for each arch, because it makes the mode way more clear, despite being longer

#if defined(__BYTE_ORDER__) && __BYTE_ORDER__ == __ORDER_LITTLE_ENDIAN__ && defined(__AVX2__)
// AVX2 version
inline void
AndPackedBytes(uint8_t* target, const uint8_t* source, size_t byte_count, unsigned shift) {...}
#elif #if defined(__BYTE_ORDER__) && __BYTE_ORDER__ == __ORDER_LITTLE_ENDIAN__ && defined(__SSE2__)
// SSE2 version
inline void
AndPackedBytes(uint8_t* target, const uint8_t* source, size_t byte_count, unsigned shift) {...}
...
// other versions
...
#else
// baseline version here
...
#endif

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

got it. I'll make the relevant changes along with it next time when there are related updates.

@alexanderguzhva

Copy link
Copy Markdown
Collaborator

/lgtm
/approve

@sre-ci-robot

Copy link
Copy Markdown
Collaborator

[APPROVALNOTIFIER] This PR is APPROVED

This pull-request has been approved by: alexanderguzhva, marcelo-cjl

The full list of commands accepted by this bot can be found here.

The pull request process is described here

Details Needs approval from an approver in each of these files:
  • OWNERS [alexanderguzhva,marcelo-cjl]

Approvers can indicate their approval by writing /approve in a comment
Approvers can cancel approval by writing /approve cancel in a comment

@alexanderguzhva

Copy link
Copy Markdown
Collaborator

issue: #1860

@alexanderguzhva

Copy link
Copy Markdown
Collaborator

/kind improvement

@sre-ci-robot
sre-ci-robot merged commit 464c368 into zilliztech:main Sep 29, 2026
13 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants