Skip to content

Inline (?i) degenerates the candidate set to the full corpus: ~1000x slower, and the summary still claims (via server) #162

Description

@Latinx

Summary

An inline (?i) is honoured by the matcher but never reaches the query planner, so trigram extraction yields no usable postings and the candidate set degenerates to the entire corpus. The search still returns correct results and still prints the (via server) summary, so the ~1000x penalty is invisible.

tgrep -c  '(?i)shellshock|bashdoor' .   # 11.1 s
tgrep -c -i 'shellshock|bashdoor' .     #  9.4 ms   (same 62 matches)

Reproduction

Any indexed repo large enough that a full scan is expensive. On a 395,558-file tree (a corpus of JSON files), tgrep 1.0.9, Linux x86_64 — same root, same server, same index:

query matches wall summary line
-P '(?i)shellshock|bashdoor' 62 11,070.9 ms (via server)
-i -P 'shellshock|bashdoor' 62 9.4 ms (via server)
'shellshock|bashdoor' 24 8.8 ms (via server)
-E utf-8 shellshock (documented index bypass) 24 4,936.7 ms Brute-force search completed in … (395558 files)

The cold-cache end of the range for the (?i) query is ~58 s.

Evidence

--trace shows the planner sees case_insensitive=false for the inline flag while the matcher applies it, and that the candidate set is the whole corpus:

[trace] search: pattern="(?i)shellshock|bashdoor" case_insensitive=false raw_candidates=395558 candidates=395558 matches=62 elapsed=57999.3ms (index=174.7ms resolve=56344.5ms search=1479.5ms)
[trace] search: pattern="shellshock|bashdoor"     case_insensitive=true  raw_candidates=6      candidates=6      matches=62 elapsed=51.2ms   (index=49.4ms resolve=0.3ms search=0.9ms)

The index is queried and returns quickly (index=174.7 ms) — it simply returns every file. The cost lands in resolve (56.3 s of 58.0 s): per-file open/read overhead, which the 50,000-entry content cache cannot absorb when the candidate set is 395,558 files. Matching itself is only 1.5 s.

That makes this a different code path from the documented index-bypassing flags (-E, -a, --binary, ignore-disabling), which do report Brute-force search completed in … (N files). Here nothing is reported, and the summary prints (via server) — the same label a 9.4 ms index-accelerated query prints.

Suggested fixes

  1. Honour the inline flag in the planner. Parse a leading (?i) and set the same internal state that the -i/--ignore-case CLI option sets, so the case-folded trigram path is used. -i already does this correctly and is 9.4 ms on the same query.
  2. Surface the degenerate case. The condition is already computed for --trace; key a warning on raw_candidates == total_files and either annotate the summary (e.g. (via server; no index narrowing)) or emit the existing Brute-force … line. This is the same reporting gap --stats should report at the end #145 asked about for --stats.

Environment

  • tgrep 1.0.9, Linux x86_64 (static-pie), index pre-built, Indexing: complete, watcher active, Cache: 50000/50000.
  • Reproduces across server restarts (fresh PID, cold cache), so it is not a stale-index artifact.
  • Absolute timings vary with content-cache state (58 s cold → ~6 s fully warm); the ratio to the -i form is ~1000x throughout.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions