Skip to content

perf: skip AST parsing without soft keywords - #2240

Merged
nedbat merged 4 commits into
coveragepy:mainfrom
KRRT7:perf/phystokens-fast-path
Jul 30, 2026
Merged

nedbat merged 4 commits into
coveragepy:mainfrom
KRRT7:perf/phystokens-fast-path

Conversation

@KRRT7

@KRRT7 KRRT7 commented Jul 21, 2026

Copy link
Copy Markdown
Contributor

Summary

Avoid parsing the full AST when looking for soft-keyword lines if the source contains no possible soft keywords.

The fast path checks for match, case, type, and lazy before calling ast.parse, while preserving the existing version-specific handling for each keyword.

Impact

Most Python files do not use soft keywords. Skipping the AST parse for those files reduces tokenization overhead during reporting.

How this optimization was found

While reviewing the physical-tokenization path, we found that find_soft_key_lines() parsed every source file even when the file could not contain a soft keyword. A cheap source-text check can identify that common case before doing any AST work.

@nedbat

nedbat commented Jul 30, 2026

Copy link
Copy Markdown
Member

I changed the match/case check so that it would only parse the file if both "match" and "case" were found in the file. A check across some random source trees:

162711 .py files
81574 use "or": 50% will be parsed
95486 use "and", 59% will be parsed

@nedbat

nedbat commented Jul 30, 2026

Copy link
Copy Markdown
Member

BTW, I also tried a regex check for complete words, but it only reduced to 48% of files being parsed.

re.search(r"\b(match|case|type|lazy)\b", source)

@nedbat
nedbat merged commit c83a0e0 into coveragepy:main Jul 30, 2026
72 checks passed
@nedbat

nedbat commented Aug 2, 2026

Copy link
Copy Markdown
Member

This is now released as part of coverage 7.15.3.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants