Repository navigation
fread "segfault from C stack overflow" using 1M+ columns #1967
Description
Activity
Reproducible on 1.10.1 and Ubuntu. Not sure why you were not able to reproduce on linux, did you try the same version?
Looks like the issue is related to amount of columns present in your file.- changed the title
[-]fwrite "segfault from C stack overflow" on MAC using big data[/-][+]fread "segfault from C stack overflow" using 1M+ columns[/+]on Dec 19, 2016 Hi @jangorecki, thank you for your reply. I think I found the problem; Is not related on OS or library version (I've tried in Linux and Windows with the last library version 1.10.0) The main point could be
stack_sizein R environment. For example, in (my) windows environment is:> Cstack_info() size current direction eval_depth 19922944 15128 1 2
While on my Debian is:
> Cstack_info() size current direction eval_depth 7969177 8848 1 2
And is a little lower than the amount needed to handle the example file. In linux (but I think also in OSX, @malbianco could you try it?) you have to increase size with
ulimit(before starting R). You need executeulimit -ato determine if size is specified inKborbytes. Then you could increase stack size (in your current terminal) with:$ ulimit -s 65535To increase size stack to 64Mb (the amount in expressed in
Kbin my linux distribution). Now, by entering R, you should see a different stack size:> Cstack_info() size current direction eval_depth 63753420 12992 1 2
Now you can load the file provide in the example. More info on R memory could be found here.
Cheers,
Paolo
@bunop, unfortunately that solution does not work for me. I am using exactly the same type of files as @malbianco.
I am able to change the stack size to 16MB, but was unable to do so for 64MB:
bash: ulimit: stack size: cannot modify limit: Operation not permittedBut even when it was changed to 16MB, Rstudio did not recognize the change when I tried
Cstack_info()...@malbianco, if in the meantime you want a temporary fix, you can read a PED file with this transposable function:
http://stackoverflow.com/questions/17288197/reading-a-csv-file-organized-horizontallyIt is faster then the normal
read.csvorread.table, but stillfreadis faster than anything else.Hi @danimag , did you get the same error type (Error: segfault from C stack overflow)? you haven't modified the stack size: see your error message: 'Operation not permitted': You should see the change with
ulimit -a(from terminal) andCstack_info()(from R). There's no need of root privileges to decrease a resource withulimit, but you need it to increase a resource. Try to do the same command withsudo. You won't see any error message. Consider thatulimitwill change limits on current terminal, so if you open another session (or if you launch Rstudio by clicking the icon on your desktop) you shouldn't see the changes. You need to open a R session (or Rstudio) from the same terminal modified withulimit.
You can also try to change limits permanently (but I suggest you to change them when you needed it) but you need to reboot your machine to see the effect. Take a look hereCheers
@bunop, I use Rstudio in Linux, so did not see any error description whatsoever.
I was able to increase the stack size to 16MB without 'sudo', but after that it didn't allow me to change it another time (in the same terminal). So I see I should have opened Rstudio through that terminal, might try that later. But that being a temporary solution, I won't be able to include it on my script, which I plan to release in the future... Or add your link to the manual.I solved the problem "temporarily". My stack size was:
stack size (kbytes, -s) 8192
using.
$ ulimit -s 16384
my stack size has increased. But this is a temporary solution and I do not know the maximum limit of my stack size. Rstudio continues to always have the same problem.
If I need open a very large file I must use the terminal, use theulimitcommand and after, run the R script.
The only solution is to make permanent the stack size value to the entire system.One could possibly include a "system" command inside Rstudio to change the stack when desired, but it might not work since you have to restart it. So for me this option does not seem viable, unfortunately, which might be the only one.
@danimag, you can't change this value with a
system(system2) command, because when this function is called, a new shell environment is created, then the command is executed and finally the terminal is closed and all its variables are lost. This environment is different from Rstudio, so you can't change it, even by calling a system and put a very long process in background. The only way is by changing it permanently (but unfortunately I don't know how to change it on Mac OSX) or temporarly (usingulimit). Then allRandRstudiosession need to start from the same terminal (I found here a tip on how to launch Rstudio from a terminal on OSX, but I have no OSX so I have no idea if this trick works; if you want to try it on linux, simply typerstudioin terminal). IMHO is better to useRterminal in dedicated infrastructures when dealing with big data.@bunop That is a good point, I did not realise that about
system. Opening R or Rstudio from the terminal in linux is just like you said, so no issues there. I would gladly do that in-house, but I don't feel like seeing that as an option if I eventually publish this script. If I do maybe I could put an optional argument to whether usedata.tablefor this or not though.I understand. My interest in this discussion is because I didn't found a solution to this problem. The main point is that such problem is not related to
data.tableitself but on your system. I realize also that such solutions are not familiar to the basic user, and need to be avoided if possible. Take a look also at bigtable (you can find an intro here). In our case this solution doesn't work because our ped files may contain letters and this is not supported at the moment by the library. Moreover, all data need to be of the same type (like a matrix). Good luck!Letters is all my ped files are made of, but thanks for the suggestion! :) It seems there are a few great alternatives to
r-basefor large datasets.I encountered the same problem. I summarize the issue and include some additional details for future reference.
There are two main cases:
- File has not a lot of columns
- File has indeed a lot of columns
File has not a lot of columns
Sometimes, there are files which have a mixed Windows and Unix line endings. I believe
fread()should fix and warn about such cases.In my case, only the first line with the headers in the
.csvhad a\r\n (CRLF) in that order, the Windows standard.fread()was confused and mistakenly thought that my large file had ~300 millions of columns, since it thought that every comma-separated value was a column! To check if this is indeed the problem, enable the-Aoption incatand you will see that the first line ends with a^M:cat -A first_three_lines.csv Timestamp,Open,High,Low,Close,Volume,Volume_(Currency),Weighted_Price^M$ 1325192180,3.237,3.237,3.237,3.237,0.3,1.6988,3.237$ 1326100360,3.1,3.1,3.1,3.1,0.623628,2.6668738,3.1$
To solve this issue, delete the
^Mcharacters manually with an editor likevimorsedor use thedos2unixutility (it should be available in your linux repositories),File has indeed a lot of columns
In the case that the file has a lot of columns (think of millions+), in linux environments, you must increase the
stack sizelimit of your session. This is done (read reference for potential dangers) either- via the shell or
- setting them via PAM modules (recommended).
You cannot set soft limits via the shell if they are larger than the hard limits.
The current
stack sizelimit is shown runningulimit -s(orulimit -a | grep 'stack size').
The soft limits (can temporarily be increased) cannot be larger than the hard limits.
Setting a limit larger than the hard limit might be the cause of thebash: ulimit: stack size: cannot modify limit: Operation not permittederror when setting a very large limit.One other cause of this error is that, in my system, I cannot increase the limit after decreasing it:
$ ulimit -s 102400 # the current stack size limit ulimit -s 20000 # decrease to 20000 kbytes $ ulimit -s 20000 # the current decreased stack size limit $ ulimit -s 21000 # bash: ulimit: stack size: cannot modify limit: Operation not permitted # cannot be increased! you have to begin a new session; open a new shell
Increase the limits with PAM
To increase the limits permanently in your system for every new session:
-
Add the
pam_limit.somodule in/etc/pam.d/loginto enable it:cat /etc/pam.d/login # <truncated output> # session required /lib/security/pam_limits.so -
Increase the hard or soft
stack sizelimit in/etc/security/limits.conf(the file itself should be well-documentated) by adding these lines:your_username hard stack 102400 your_username soft stack 8192 # or more if you never want to use ulimit -s -
Reboot or reload/restart your PAM configuration/service.
Note that even with a large limit
Cstack_info()can reportNA(I do not know why), without limitingfread().ulimit -s 102300; ulimit -s; Rscript <(echo 'Cstack_info()') # 102300 # The stack size limit is 102300 # size current direction eval_depth # NA NA 1 2 # stack size limit is 102300 but R reports NA ulimit -s 82300; ulimit -s; Rscript <(echo 'Cstack_info()') # 82300 # The stack size limit is 82300 # size current direction eval_depth # 80061440 8320 1 2 # stack size limit is 82300 and R does not report NA
Setting the limit appropriately, you should be able to read the file with a very large number of columns.
In my case, setting a limit withulimit -s 102300and no less, I could read a file with around 13,000,000 columns!Errors that R/data.table may throw:
-
When the required
stack sizelimit is very close to the required limit by the number of the millions of columns, I got this error:Error: C stack usage 8036772 is too close to the limit -
When the number of columns were quite larger for the required
stack sizelimit, I got:Error: segfault from C stack overflow -
R crashed when the number of columns was even larger!:
Segmentation fault (core dumped)or
*** caught segfault *** address 0x7fffce96a820, cause 'memory not mapped' Traceback: 1: fread("data.csv") Possible actions: <trunccated>
Reacted by Paolo Cozzi and Pasha StetsenkoThe advice on how to increase stack size is very helpful; however latest version of
freaddoes not attempt to allocate on the stack any more, so hopefully these tweaks should not be necessary. In particular, on my system with stack size 8192KB the example file can be read without any hiccups.The problem of inconsistent line endings is a separate problem, and deserves a dedicated issue.
Hi, While the solution works for R. It doesn't appear to increase stack size for Rstudio-server.
Could you provide any advise on how to get the stack increased for rstudio-serverec2 [ami instance , in particular]?Hi. The point is that changing the ulimit value in a terminal affect only R instances in that terminal, as described in the discussion. This change isn't permanent, so if you relog in a different terminal or reboot the machine, you have to apply the same command. Maybe rstudio server starts from a init script, and so ulimit has no effect on it. if you have a rstudio server instance on a machine (regardless if it is a physical, virtual, o even in a docker container) the easiest solution could be change this value system wide or set this value in the same (init) script which starts Rstudio server, or start rstudio server manually in the same unlimited terminal. Hope this helps
I tried to. I did set the value in a permanent way. But it wont help get Cstack_info() in Rstudio-server changed. However, it works brilliantly with R.
I as well approached to stop rstudio-server and start it from /sbin/rstudio-server from the terminal. It wont help,m unfortunately.
How to "set this value in the same (init) script which starts Rstudio server" ?
I tried to execute system(ulimit -s 65535) under rstudio-server web interface neither it impacts things, as it seems to meI think if you change this values system wide (and you see this value changed in each of your terminal) but rstudio server still report the same value could be that rstudio server has its ulimit configuration file or you have changed the ulimit value for a user that isn't the user which starts rstudio. If it is a issue regarding rstudio, maybe post this problem into its community may helps. You can't change this value using a system in rstudio server, I've explained why before in the discussion
I tried both. I posted to community forum. Bur rstudio pro [tested with amazon ec2] wont support the Cstack increase as well, as it seems to me. I posted to Rstudio Pro forum. hopefully they will respond with some meaningful advise soon.
References:
https://support.rstudio.com/hc/en-us/requests/24888
https://support.rstudio.com/hc/en-us/community/posts/115008248313-Cstack-info-value-wont-change-in-rstudio-server
I'm using "fread" function to read a file with 160 rows and columns 1141430.
The function will stop reporting an error "Error: segfault from C stack overflow" and closing the R session.
The file contains letters. Only the first two columns contain names.
This problem occurs only in the Mac system, and works well on Linux and Windows environments.
Mac OSX: 10.12.2
Version: 1.10.0 data.table
Files: http://datadryad.org/bitstream/handle/10255/dryad.104088/Dryad_Submission.7z?sequence=1
Info files: http://datadryad.org/resource/doi:10.5061/dryad.1p7sf/1
R script:
genotype_path='/Users/Gabriele/Desktop/Dryad_Submission-0/4H_160indivs_Final.ped'
pops <- data.table::fread(genotype_path, sep = " ", header = FALSE, verbose=TRUE)
verbose output :
Input contains no \n. Taking this to be a filename to open
File opened, filesize is 0.340174 GB.
Memory mapping ... ok
Detected eol as \r\n (CRLF) in that order, the Windows standard.
Positioned on line 1 after skip or autostart
This line is the autostart and not blank so searching up for the last non-blank ... line 1
Using supplied sep ' ' ... found ok
Detected 1141430 columns. Longest stretch was from line 1 to line 30
Starting data input on line 1 (either column names or first row of data). First 10 characters: Jacobs H70
'header' changed by user from 'auto' to FALSE
Count of eol: 161 (including 1 at the end)
Count of sep: 182628640
nrow = MIN( nsep [182628640] / (ncol [1141430] -1), neol [161] - endblanks [1] ) = 160
Type codes (point 0): 44111144444444444444444444444444444444444444404444444444444444444444444444444444444444444444444444444444444444444444444444444444444444444444444444444444444444444444444444444444444444444444444444000044444444440044444444000044444444444444444444444444004444444444444444444444444444444444004400440044444444444444444444444444004444444444444400444444444444444444444444444444444444444444444400444400440044444444444444444444444444444444444444444444444400444400444444444444444444444444444444444444444444004400440044004400444444444444444444444444444444440000004444004400444444444444440044000044444444444444444444444444440044004400444400444444440000444444444444444444444444440044444444444444444444444444444444004444440044004444444400444444444400444444004444444444004444004400444444444444444444004444004444444400444444444444444444440044444444004444444444444444444444000044000044444400440044444444444444444444004444444444444444444444444444444444444444444444444400444444444400444444004444444...
Could you help me with this?