Skip to content

perf(core): share frizbee hot-loop instantiations to cut binary size ~17% - #872

Merged
dmtrKovalenko merged 4 commits into
mainfrom
chore/optimize-binary-size
Sep 21, 2026
Merged

dmtrKovalenko merged 4 commits into
mainfrom
chore/optimize-binary-size

Conversation

@dmtrKovalenko

@dmtrKovalenko dmtrKovalenko commented Sep 15, 2026 •

Copy link
Copy Markdown
Owner

Why

libfff_nvim.so is ~15MB. Almost 70% of its code is neo_frizbee: the matcher hot loop is monomorphized per call site × 8 SIMD backends × 10 typo/unicode variants. Three of the six call sites existed only because fff passed distinct haystack types (FileItem, &FileItem, String) for what is the same loop.

What

  • match_fuzzy_parts always feeds frizbee &[&FileItem]. The unconstrained "all files" path now collects a slice of references (one pointer per file) instead of carrying its own FileItem instantiation.
  • fuzzy_match_byte_offsets_for_page matches on &str instead of String, sharing the instantiation already used by fuzzy grep.

Result (--profile ci, x86_64 linux)

before after
libfff_nvim.so 14.9 MB 12.4 MB
.text 11.5 MB 9.1 MB
frizbee symbols 607 454

No behavior change. make lint-rust and make test-rust pass.

Follow-ups (frizbee side)

Remaining frizbee code is still ~7MB. Unicode variants are 3.3× the ASCII ones, and the scalar fallback is the single largest backend (2.5MB) while only reachable on x86_64 without SSE4.1. Those belong in the frizbee repo.

Summary by CodeRabbit

  • Improvements
    • Improved consistency and reliability of fuzzy matching across files, directories, pages, and filtered results.
    • Enhanced fuzzy searches to better handle short queries and minor typos.
    • Improved matching of file paths and filenames, including fallback behavior.
    • Increased compatibility with updated fuzzy-search behavior while preserving unsorted result handling.
    • Improved result ranking by incorporating Git recency, including when recency-based filtering is used.

@coderabbitai

coderabbitai Bot commented Sep 15, 2026 •

Copy link
Copy Markdown

Review Change StackReview Change Stack

Understand this PR’s impact

Explore downstream dependencies and potential security impact with Blast Radius.

View blast radius →

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Advanced

Run ID: cdc6c4d2-32af-48de-80d7-9d6d19ce9f8b

📥 Commits

Reviewing files that changed from the base of the PR and between 3f3e9f3 and af2a2c3.

⛔ Files ignored due to path filters (1)
  • Cargo.lock is excluded by !**/*.lock
📒 Files selected for processing (2)
  • Cargo.toml
  • crates/fff-core/src/score.rs

Included review availability: Your plan provides up to 10 included reviews per hour; 9 remain after this review.


📝 Walkthrough

Walkthrough

The change upgrades neo_frizbee and updates matcher configuration, matcher calls, character-index conversions, and fuzzy-match span handling.

Changes

neo_frizbee migration

Layer / File(s) Summary
Dependency and matcher API updates
Cargo.toml, crates/fff-core/src/grep/fuzzy_grep.rs, crates/fff-core/src/grep/grep_tests.rs
The workspace upgrades neo_frizbee to 0.13.3. Matching code and tests use SortStrategy::Unsorted and the Matcher API. Release builds strip symbols.
Character-index compatibility
crates/fff-core/src/grep/sink.rs, crates/fff-core/src/grep/fuzzy_grep.rs, crates/fff-core/src/grep/grep_tests.rs
Character indices use u32 at the matcher boundary and convert to usize for indexing. Fuzzy-match span calculations cast explicitly to usize.

Priority: ⬇️ Low

Estimated code review effort: 2 (Simple) | ~12 minutes

Change: Refactor

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 76.92% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 13 functions across 4 files. (1 skipped: … Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately identifies the main change: sharing frizbee hot-loop instantiations to reduce core binary size. It is concise and specific.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Full details: Docstring Coverage

Explanation

Docstring coverage is 76.92% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 13 functions across 4 files. (1 skipped: 1 unsupported.)

  • Fix all pre-merge checks with AI
✨ Finishing Touches 💡 1
📝 Generate docstrings 💡
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@greptile-apps

greptile-apps Bot commented Sep 15, 2026

Copy link
Copy Markdown

Confidence Score: 4/5

The PR appears safe to merge, though the new full-index allocation on every unconstrained search warrants performance measurement for large repositories.

Borrowing FileItem and string inputs preserves their ordering and contents, leaving no established correctness failure; the remaining concern is non-blocking runtime overhead introduced on the common search path.

Files Needing Attention: crates/fff-core/src/score.rs

Important Files Changed

Filename Overview
crates/fff-core/src/score.rs Unifies matcher input types to reduce binary size, with a non-blocking concern about the new per-search O(n) reference-vector allocation.

Flowchart

%%{init: {'theme': 'neutral'}}%%
flowchart LR
    A[FileItems::All] --> B[Collect Vec of FileItem references]
    C[FileItems::Filtered] --> D[Existing reference slice]
    B --> E[Shared frizbee FileItem-reference instantiation]
    D --> E
    F[Page path Strings] --> G[Create str views]
    G --> H[Shared frizbee str instantiation]
Loading

Reviews (1): Last reviewed commit: "perf(core): share frizbee hot-loop insta..." | Re-trigger Greptile

Comment thread crates/fff-core/src/score.rs Outdated
let all_refs: Vec<&FileItem>;
let candidates: &[&FileItem] = match working_files {
FileItems::All(files) => {
all_refs = files.iter().collect();

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Full-index allocation per search

The common unconstrained-search path now builds a fresh Vec<&FileItem> containing every indexed file before each match. The previous implementation passed the existing file slice directly. For large indexes, repeated searches therefore add O(n) work and temporary memory use before matching starts. This is non-blocking, but the hot-path cost should be benchmarked and, if material, avoided with reusable storage or a design that does not rebuild the reference vector.

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

- rebased fork on upstream frizbee 0.13.0; drops the scalar SIMD fallbacks that
  can never be selected on aarch64 (~3 MB of inlined dead dispatch code) and adds
  `SortStrategy::Unsorted`, which skips the parallel k-merge fff never needed
- API migration: `Config.sort` is now an enum, match indices are `Vec<u32>`, and
  the free `match_list*` functions became `Matcher::new(..).match_list*(..)`

| `libfff_nvim.dylib` (release, aarch64) | before | after |
|---|---|---|
| file size | 8.19 MB | 6.76 MB |
| `__text` | 5.85 MB | 4.45 MB |

Search latency is unchanged within noise (+1 to +3% on fff's fuzzy_search bench,
identical results).
- `neo_frizbee` 0.13.1 adds `match_range_parallel_resolved(len, resolve, threads)`,
  resolving items by index instead of through a `&[T]` slice
- drops the `Vec<&FileItem>` built for the full list and the per-part `subset`
  vectors for files and dirs; each pass now addresses survivors through the
  previous round's matches directly

```rust
neo_frizbee::match_range_parallel_resolved(
    part,
    survivors.len(),
    &|index, buf: &mut [*const u8; MAX_PATH_CHUNKS]| {
        let file = working_files.index(survivors[index as usize].index as usize);
        resolve_file_chunks(file, arena, buf)
    },
    &part_options,
    max_threads,
);
```
@dmtrKovalenko
dmtrKovalenko force-pushed the chore/optimize-binary-size branch from 3f3e9f3 to af2a2c3 Compare September 21, 2026 02:46
@dmtrKovalenko
dmtrKovalenko merged commit 246683b into main Sep 21, 2026
53 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant