Skip to content

Faster codemirror search match counting - #612

Merged
gschier merged 1 commit into
mainfrom
fix/search-match-count-perf
Aug 31, 2026
Merged

gschier merged 1 commit into
mainfrom
fix/search-match-count-perf

Conversation

@gschier

@gschier gschier commented Aug 31, 2026 •

Copy link
Copy Markdown
Member

Summary

Typing into the response search panel (Cmd+F) lags badly on multi-megabyte bodies. Fixes Improve search performance for large responses.

The panel's own search is viewport-bounded and fine. The match count beside it isn't: searchMatchCount.ts ran SearchCursor over the entire document on every query change and every selection change — ~400ms per keystroke on a 5MB response, and the 9999-match cap stopped helping as soon as the query got specific enough to have few matches. The JSONPath filter field is not involved; it applies on Enter.

Where the time actually goes

Measured on 5M characters:

operation time
doc.toString() 3.3ms
chunked indexOf over doc.iter() 2.4ms
text.normalize("NFKD") on the whole string 4.3ms
normalize("NFKD") once per code point — what SearchCursor does 333ms
SearchCursor, whole document 365ms

91% of the cursor's time is one ICU call per code point. The same Unicode work in bulk is 77× cheaper, so no amount of caching, capping or debouncing makes a cursor-based count fast. It has to stop normalizing a character at a time. (@codemirror/search 6.7.1 has no match-count API to lean on — it exports the two cursors and SearchQuery, and the discuss thread is people writing this loop themselves.)

Implementation

normalizeDoc rewrites the document the way the cursor does — NFKD, then a case fold unless the search is case-sensitive — but once per document instead of once per character. Each distinct character is normalized once and the answer reused; runs of ASCII never reach ICU. Doing it per character rather than over the whole string still matters: it stops NFKD from reordering combining marks across characters, and stops toLowerCase from applying Greek final-sigma rules the cursor doesn't.

Characters whose normalized form has a different length are recorded as expansions, so matches found with indexOf map back to document offsets — floor for the start, ceil for the end, which is what the cursor reports when a match begins or ends inside a decomposed character.

Matches are kept per document and query, so moving the selection looks up the current one instead of scanning again. Regexp and whole-word queries still go through the cursor.

Measurements

Synthetic 5MB JSON, time to count one query:

document query before after
ASCII user123 427ms 3.2ms
CJK + emoji user123 415ms 2.5ms
café in every row user123 402ms 2.8ms
café in every row café 400ms 1.6ms

Plus a one-time normalized copy on the first keystroke: 17ms (ASCII), 17ms (accented), 40ms (CJK-heavy).

Tests

The offset mapping is the part that can go quietly wrong, and only on input nobody thinks to write a case for — so the input is generated. A seeded fuzz test builds documents from an alphabet of precomposed and decomposed accents, ligatures, ellipses, NBSP, İ, Σ/ς, CJK, emoji, roman numerals, fullwidth letters, bare combining marks and backslashes, then asserts the counter's matches equal SearchCursor's exactly, from and to, in both case modes.

It earned its keep immediately — it found both correctness rules in the diff, neither of which I would have reasoned my way to:

  • a match ends on a whole code point however the query was cut (a lone surrogate in the query)
  • scanning resumes past the character a match ended inside, so . counts … once rather than three times

Before merging I ran it at 20,000 rounds × 5 seeds × 2 case modes — 200,000 generated documents, exact agreement throughout. It's committed at 400 rounds so it stays cheap in CI.

Alongside that: overlap skipping, \n unquoting vs literal, the 9999 cap, current-match lookup, the regexp/whole-word handover, and match reuse across a selection-only transaction. vp test 519/519, tsc --noEmit and oxlint clean.

Notes

Worth reporting upstream — SearchCursor could normalize per chunk, or skip it when both sides are ASCII, and every CodeMirror user counting matches would get this for free.

🤖 Generated with Claude Code

@greptile-apps

greptile-apps Bot commented Aug 31, 2026 •

Copy link
Copy Markdown

Greptile Summary

The PR optimizes response-search match counting by caching normalized document text and match ranges, while retaining CodeMirror’s cursor for complex queries.

  • Adds a plain-query indexOf counting path with Unicode-aware normalization and document-offset mapping.
  • Reuses cached matches for selection-only updates.
  • Adds deterministic and generated tests comparing match ranges against CodeMirror’s search cursor.

Confidence Score: 5/5

The PR appears safe to merge.

No blocking failure remains.

Important Files Changed

Filename Overview
apps/yaak-client/components/core/Editor/searchMatchCount.ts Adds cached match counting, a fast path for plain searches, Unicode normalization, and current-match lookup without an eligible follow-up defect.
apps/yaak-client/components/core/Editor/searchMatchCount.test.ts Adds broad parity coverage for counting, normalization, offsets, cache reuse, and generated Unicode inputs.

Reviews (2): Last reviewed commit: "fix(editor): count search matches withou..." | Re-trigger Greptile

greptile-apps[bot]
greptile-apps Bot previously approved these changes Aug 31, 2026
…a time

Typing into the response search panel lagged badly on multi-megabyte
bodies. The panel's own search is viewport-bounded, but the match count
next to it isn't: it ran SearchCursor over the entire document on every
query change and every selection change, around 400ms per keystroke on a
5MB response.

91% of that is one ICU normalize() call per code point. Normalizing the
same five million characters in one call takes 4ms, so no amount of
caching or scheduling makes a cursor-based count fast — it has to stop
normalizing a character at a time.

The count now rewrites the document the way the cursor does, NFKD and
then a case fold, once per document rather than once per character:
each distinct character is normalized once and the answer reused, and
runs of ASCII never reach ICU at all. Characters whose normalized form
is a different length are recorded, so matches found with indexOf can be
reported back in document offsets. Counting is then ~3ms for any query
on any document, and the matches are kept so moving the selection looks
up the current one instead of scanning again.

Two rules keep the results identical to the cursor's: a match ends on a
whole code point however the query was cut, and scanning resumes past
the character a match ended inside, so a query of "." counts "…" once
rather than three times. Both came out of fuzzing against the cursor
rather than from reasoning about it, so that test stays.

Regexp and whole-word queries still go through the cursor.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@gschier
gschier force-pushed the fix/search-match-count-perf branch from 9b17b89 to fe09181 Compare August 31, 2026 16:52
@greptile-apps
greptile-apps Bot dismissed their stale review August 31, 2026 16:52

Dismissed because a newer commit was pushed; Greptile will re-review the current head.

@gschier gschier changed the title fix(editor): stop the search match count from rescanning the whole response per keystroke fix(editor): count search matches without normalizing a character at a time Aug 31, 2026
@gschier gschier changed the title fix(editor): count search matches without normalizing a character at a time Faster codemirror search match counting Aug 31, 2026
@gschier
gschier merged commit 661a384 into main Aug 31, 2026
8 checks passed
@gschier
gschier deleted the fix/search-match-count-perf branch August 31, 2026 17:09
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant