You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Inline (?i) degenerates the candidate set to the full corpus: ~1000x slower, and the summary still claims (via server) #162
An inline (?i) is honoured by the matcher but never reaches the query planner, so trigram extraction yields no usable postings and the candidate set degenerates to the entire corpus. The search still returns correct results and still prints the (via server) summary, so the ~1000x penalty is invisible.
tgrep -c '(?i)shellshock|bashdoor'.# 11.1 s
tgrep -c -i 'shellshock|bashdoor'.# 9.4 ms (same 62 matches)
Reproduction
Any indexed repo large enough that a full scan is expensive. On a 395,558-file tree (a corpus of JSON files), tgrep 1.0.9, Linux x86_64 — same root, same server, same index:
query
matches
wall
summary line
-P '(?i)shellshock|bashdoor'
62
11,070.9 ms
(via server)
-i -P 'shellshock|bashdoor'
62
9.4 ms
(via server)
'shellshock|bashdoor'
24
8.8 ms
(via server)
-E utf-8 shellshock (documented index bypass)
24
4,936.7 ms
Brute-force search completed in … (395558 files)
The cold-cache end of the range for the (?i) query is ~58 s.
Evidence
--trace shows the planner sees case_insensitive=false for the inline flag while the matcher applies it, and that the candidate set is the whole corpus:
The index is queried and returns quickly (index=174.7 ms) — it simply returns every file. The cost lands in resolve (56.3 s of 58.0 s): per-file open/read overhead, which the 50,000-entry content cache cannot absorb when the candidate set is 395,558 files. Matching itself is only 1.5 s.
That makes this a different code path from the documented index-bypassing flags (-E, -a, --binary, ignore-disabling), which do report Brute-force search completed in … (N files). Here nothing is reported, and the summary prints (via server) — the same label a 9.4 ms index-accelerated query prints.
Suggested fixes
Honour the inline flag in the planner. Parse a leading (?i) and set the same internal state that the -i/--ignore-case CLI option sets, so the case-folded trigram path is used. -i already does this correctly and is 9.4 ms on the same query.
Surface the degenerate case. The condition is already computed for --trace; key a warning on raw_candidates == total_files and either annotate the summary (e.g. (via server; no index narrowing)) or emit the existing Brute-force … line. This is the same reporting gap --stats should report at the end #145 asked about for --stats.
Environment
tgrep 1.0.9, Linux x86_64 (static-pie), index pre-built, Indexing: complete, watcher active, Cache: 50000/50000.
Reproduces across server restarts (fresh PID, cold cache), so it is not a stale-index artifact.
Absolute timings vary with content-cache state (58 s cold → ~6 s fully warm); the ratio to the -i form is ~1000x throughout.
Summary
An inline
(?i)is honoured by the matcher but never reaches the query planner, so trigram extraction yields no usable postings and the candidate set degenerates to the entire corpus. The search still returns correct results and still prints the(via server)summary, so the ~1000x penalty is invisible.Reproduction
Any indexed repo large enough that a full scan is expensive. On a 395,558-file tree (a corpus of JSON files), tgrep 1.0.9, Linux x86_64 — same root, same server, same index:
-P '(?i)shellshock|bashdoor'(via server)-i -P 'shellshock|bashdoor'(via server)'shellshock|bashdoor'(via server)-E utf-8 shellshock(documented index bypass)Brute-force search completed in … (395558 files)The cold-cache end of the range for the
(?i)query is ~58 s.Evidence
--traceshows the planner seescase_insensitive=falsefor the inline flag while the matcher applies it, and that the candidate set is the whole corpus:The index is queried and returns quickly (
index=174.7 ms) — it simply returns every file. The cost lands inresolve(56.3 s of 58.0 s): per-file open/read overhead, which the 50,000-entry content cache cannot absorb when the candidate set is 395,558 files. Matching itself is only 1.5 s.That makes this a different code path from the documented index-bypassing flags (
-E,-a,--binary, ignore-disabling), which do reportBrute-force search completed in … (N files). Here nothing is reported, and the summary prints(via server)— the same label a 9.4 ms index-accelerated query prints.Suggested fixes
(?i)and set the same internal state that the-i/--ignore-caseCLI option sets, so the case-folded trigram path is used.-ialready does this correctly and is 9.4 ms on the same query.--trace; key a warning onraw_candidates == total_filesand either annotate the summary (e.g.(via server; no index narrowing)) or emit the existingBrute-force …line. This is the same reporting gap--statsshould report at the end #145 asked about for--stats.Environment
Indexing: complete, watcher active,Cache: 50000/50000.-iform is ~1000x throughout.