Repository navigation
fread spends too much time in is_url/is_secureurl/is_file for long in memory input #2531
Description
Activity
Seems like a problem of the regex engine more than anything. anchoring to string beginning should mean that
greplperformance is O(1), as far as I'm concerned. Strange that it appears otherwise.probably R's regex engine attempts to build a histogram of the target string before applying the regexp.
i have in mind this:
Indeed it may be more a R issue, but something one can at least work around in
freadUsing
perl=TRUEdoes not seem to add any negative aspect and will be usally faster (in my benchmarks 100 times for large input)See also help-page:
?grep:If you are doing a lot of regular expression matching, including on very long strings, you will want to consider the options used. Generally PCRE will be faster than the default regular expression engine,
The additional penalty (even when using
perl=TRUE) does not seem to come fromgrep: Look at the benchmark when feeding grep with the shortened string thoughsubstr(grep_perl_substr) whereas usingstringi::stri_subperforms with O(1) (grep_perl_strisub). So there seems to be an overhead when calling the base function with a large string. (Actually we see this even for the Rcpp versions)For later reference, the utility
startsWithis easily the fastest (and clearest), but unfortunately depends onR (>= 3.3.0).is_url <- function(x) grepl("^(http|ftp)s?://", x) is_url_perl <- function(x) grepl("^(?:http|ftp)s?://", x, perl = TRUE) is_url2 <- function(x) { `||`(startsWith(x, "ht") && startsWith(x, "http://") || startsWith(x, "https://"), startsWith(x, "ft") && startsWith(x, "ftp://") || startsWith(x, "ftps://")) } microbenchmark::microbenchmark(is_url(input), is_url_perl(input), is_url2(input), times = 50) Unit: microseconds # expr min lq mean median uq max neval cld # is_url(input) 1001673.642 1038144.688 1076978.567 1070702.5405 1112537.553 1185685.65 50 c # is_url_perl(input) 23280.392 24793.807 30226.567 27345.3795 30737.237 51998.24 50 b # is_url2(input) 1.807 2.711 139.927 12.8005 18.975 6414.17 50 a URL <- "https://github.com/Rdatatable/data.table/issues/2531" microbenchmark::microbenchmark(is_url(URL), is_url_perl(URL), is_url2(URL), times = 50) # expr min lq mean median uq max neval cld # is_url(URL) 6.626 7.228 8.40290 7.680 8.132 30.118 50 b # is_url_perl(URL) 65.657 66.259 67.40972 66.560 66.862 92.462 50 c # is_url2(URL) 1.506 1.808 2.25304 2.108 2.409 8.132 50 a
Reacted by javrucebo and Matt DowleThis was a great report -- many thanks!
I've just fixed it and pushed to master. Please test and reopen if this doesn't fix it. I'm pretty sure it's much better now, but you never know. I pushed it straight to master (rather than a PR) so you can more easily test it just by upgrading to latest dev.
I added your test too so it now reads those http:// addresses as data as you suggested.- added a commit that references this issue
on Feb 16, 2018 Thank you for looking into this issue.
I am afraid it is not fixing the issue of slow execution time completely.
In the new code, once you go in the else clause ofif (!missing(file))the first thing you are checking (after verifyinginputis a length 1 character vector) isif (file.exists(input)) { ....unfortunately
file.existsis also very slow for large input:input <- paste(rep("1,2,3,4.567,some text field", 1e6),collapse="\n") system.time(file.exists(input)) ## user system elapsed ## 0.693 0.000 0.694 system.time(fread(input)) ## user system elapsed ## 0.861 0.000 0.862with the
file.existseating up 80% of the totalfreadexecution time.Fortunately the next check you are doing with
grep(detecting newlines) is very fast (opposed to the suprisingly slowgreplchecks:system.time(length(grep('\\n|\\r', input))) ## user system elapsed ## 0 0 0The easiest fix seems to be to swap the two conditions. The only case where I see this would make a difference is, if
inputis a filename and has a newline in the name, but I hope no ones would do this and then expectfreadto be able to read the file. (and actually it wouldn't work as the c-routine for fread checks again for newlines and would treat such a file name as the actual input)You suggested me to re-open the issue in case it's not working, but as reporter, I can not re-open it.
I have opened a PR #2630 with the proposed fix.
The other part of your fix with substring etc is indeed avoiding
greplis slightly faster, but due to the fact you check now for newlines earlier in the code the penalty should not be visible ... unless someone is giving as input a very long string without any newlines ... but that seems also like a pathological case.Thanks for the follow up. Yes I had thought file.exists() would be fast on huge input as the operating system would detect it was invalid in the first few characters. Interesting it isn't. Maybe R passes over it first before calling the OS for some reason. Anyway, will proceed with your PR, thanks.
Didn't know you can't reopen issues; must be just project members that see that option then. I'll see if there's a project setting to open that, and will bear it in mind in future. I've invited you to be project member too.

Summary
When
freadis fed with a character string as input the routine spends considerable amount of time detecting that the supplied input is not a filename or an url.This is due to
greplnot scaling well for large input as used in fread.Example
Although the pattern is anchored at the beginning of the string, running
greplfor large inputs will take a lot of time for large inputs (more detailed benchmarks further down).This will lead then to the full call to
freadpossibly spending a third of the time for those supposedly simple checks. (example also below)Possible solutions could be one or more of the following
greplto use PERL regexp engineperl=TRUEAlternatively use another method to determine whether the input starts with url (see benchmarks below)
str=to denote that the input is to be considered as the data and skip the tests for url or file. This would be similar in spirit to thefile=argument.As a side-effect this would also allow input to be read which consits only of url's and having no header, e.g.
Profiling example
Benchmarking
grepland friendsComparing different functions to verify whether a string starts with any of http(s)/ftp(s) or file shows that
greplscales badly and is by far the slowest of the tested variants.Adding simply
perl=TRUEalready improves by around factor 100 for large inputs (code further below)sessionInfo
Results are similar with R 3.4.3 / data.table 1.10.4 / Windows 10 64bit