Repository navigation
perf: skip AST parsing without soft keywords - #2240
Merged
Merged
Conversation
With "and", still checks 59% of files. With "or" only checks 50% of files.
Member
|
I changed the match/case check so that it would only parse the file if both "match" and "case" were found in the file. A check across some random source trees: |
Member
|
BTW, I also tried a regex check for complete words, but it only reduced to 48% of files being parsed. |
Member
|
This is now released as part of coverage 7.15.3. |
1 task
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Avoid parsing the full AST when looking for soft-keyword lines if the source contains no possible soft keywords.
The fast path checks for
match,case,type, andlazybefore callingast.parse, while preserving the existing version-specific handling for each keyword.Impact
Most Python files do not use soft keywords. Skipping the AST parse for those files reduces tokenization overhead during reporting.
How this optimization was found
While reviewing the physical-tokenization path, we found that
find_soft_key_lines()parsed every source file even when the file could not contain a soft keyword. A cheap source-text check can identify that common case before doing any AST work.