Skip to content

fread "segfault from C stack overflow" using 1M+ columns #1967

Description

@malbianco

I'm using "fread" function to read a file with 160 rows and columns 1141430.
The function will stop reporting an error "Error: segfault from C stack overflow" and closing the R session.
The file contains letters. Only the first two columns contain names.
This problem occurs only in the Mac system, and works well on Linux and Windows environments.

Mac OSX: 10.12.2
Version: 1.10.0 data.table
Files: http://datadryad.org/bitstream/handle/10255/dryad.104088/Dryad_Submission.7z?sequence=1
Info files: http://datadryad.org/resource/doi:10.5061/dryad.1p7sf/1
R script:
genotype_path='/Users/Gabriele/Desktop/Dryad_Submission-0/4H_160indivs_Final.ped'
pops <- data.table::fread(genotype_path, sep = " ", header = FALSE, verbose=TRUE)

verbose output :

Input contains no \n. Taking this to be a filename to open
File opened, filesize is 0.340174 GB.
Memory mapping ... ok
Detected eol as \r\n (CRLF) in that order, the Windows standard.
Positioned on line 1 after skip or autostart
This line is the autostart and not blank so searching up for the last non-blank ... line 1
Using supplied sep ' ' ... found ok
Detected 1141430 columns. Longest stretch was from line 1 to line 30
Starting data input on line 1 (either column names or first row of data). First 10 characters: Jacobs H70
'header' changed by user from 'auto' to FALSE
Count of eol: 161 (including 1 at the end)
Count of sep: 182628640
nrow = MIN( nsep [182628640] / (ncol [1141430] -1), neol [161] - endblanks [1] ) = 160
Type codes (point 0): 44111144444444444444444444444444444444444444404444444444444444444444444444444444444444444444444444444444444444444444444444444444444444444444444444444444444444444444444444444444444444444444444444000044444444440044444444000044444444444444444444444444004444444444444444444444444444444444004400440044444444444444444444444444004444444444444400444444444444444444444444444444444444444444444400444400440044444444444444444444444444444444444444444444444400444400444444444444444444444444444444444444444444004400440044004400444444444444444444444444444444440000004444004400444444444444440044000044444444444444444444444444440044004400444400444444440000444444444444444444444444440044444444444444444444444444444444004444440044004444444400444444444400444444004444444444004444004400444444444444444444004444004444444400444444444444444444440044444444004444444444444444444444000044000044444400440044444444444444444444004444444444444444444444444444444444444444444444444400444444444400444444004444444...

Could you help me with this?

Activity

  1. jangorecki commented on Dec 19, 2016

    @jangorecki
    Member

    Reproducible on 1.10.1 and Ubuntu. Not sure why you were not able to reproduce on linux, did you try the same version?
    Looks like the issue is related to amount of columns present in your file.

  2. changed the title [-]fwrite "segfault from C stack overflow" on MAC using big data[/-] [+]fread "segfault from C stack overflow" using 1M+ columns[/+] on Dec 19, 2016
  3. bunop commented on Dec 23, 2016

    @bunop

    Hi @jangorecki, thank you for your reply. I think I found the problem; Is not related on OS or library version (I've tried in Linux and Windows with the last library version 1.10.0) The main point could be stack_size in R environment. For example, in (my) windows environment is:

    > Cstack_info()
          size    current  direction eval_depth 
      19922944      15128          1          2 

    While on my Debian is:

    > Cstack_info()
          size    current  direction eval_depth 
       7969177       8848          1          2

    And is a little lower than the amount needed to handle the example file. In linux (but I think also in OSX, @malbianco could you try it?) you have to increase size with ulimit (before starting R). You need execute ulimit -a to determine if size is specified in Kb or bytes. Then you could increase stack size (in your current terminal) with:

    $ ulimit -s 65535

    To increase size stack to 64Mb (the amount in expressed in Kb in my linux distribution). Now, by entering R, you should see a different stack size:

    > Cstack_info()
          size    current  direction eval_depth 
      63753420      12992          1          2 

    Now you can load the file provide in the example. More info on R memory could be found here.

    Cheers,

    Paolo

  4. danimfernandes commented on Dec 26, 2016

    @danimfernandes

    @bunop, unfortunately that solution does not work for me. I am using exactly the same type of files as @malbianco.

    I am able to change the stack size to 16MB, but was unable to do so for 64MB:
    bash: ulimit: stack size: cannot modify limit: Operation not permitted

    But even when it was changed to 16MB, Rstudio did not recognize the change when I tried Cstack_info()...

  5. danimfernandes commented on Dec 26, 2016

    @danimfernandes

    @malbianco, if in the meantime you want a temporary fix, you can read a PED file with this transposable function:
    http://stackoverflow.com/questions/17288197/reading-a-csv-file-organized-horizontally

    It is faster then the normal read.csv or read.table, but still fread is faster than anything else.

  6. bunop commented on Dec 27, 2016

    @bunop

    Hi @danimag , did you get the same error type (Error: segfault from C stack overflow)? you haven't modified the stack size: see your error message: 'Operation not permitted': You should see the change with ulimit -a (from terminal) and Cstack_info() (from R). There's no need of root privileges to decrease a resource with ulimit, but you need it to increase a resource. Try to do the same command with sudo. You won't see any error message. Consider that ulimit will change limits on current terminal, so if you open another session (or if you launch Rstudio by clicking the icon on your desktop) you shouldn't see the changes. You need to open a R session (or Rstudio) from the same terminal modified with ulimit.
    You can also try to change limits permanently (but I suggest you to change them when you needed it) but you need to reboot your machine to see the effect. Take a look here

    Cheers

  7. danimfernandes commented on Dec 27, 2016

    @danimfernandes

    @bunop, I use Rstudio in Linux, so did not see any error description whatsoever.
    I was able to increase the stack size to 16MB without 'sudo', but after that it didn't allow me to change it another time (in the same terminal). So I see I should have opened Rstudio through that terminal, might try that later. But that being a temporary solution, I won't be able to include it on my script, which I plan to release in the future... Or add your link to the manual.

  8. malbianco commented on Jan 13, 2017

    @malbianco
    Author

    I solved the problem "temporarily". My stack size was:
    stack size (kbytes, -s) 8192
    using.
    $ ulimit -s 16384
    my stack size has increased. But this is a temporary solution and I do not know the maximum limit of my stack size. Rstudio continues to always have the same problem.
    If I need open a very large file I must use the terminal, use the ulimit command and after, run the R script.
    The only solution is to make permanent the stack size value to the entire system.

  9. danimfernandes commented on Jan 13, 2017

    @danimfernandes

    One could possibly include a "system" command inside Rstudio to change the stack when desired, but it might not work since you have to restart it. So for me this option does not seem viable, unfortunately, which might be the only one.

  10. bunop commented on Jan 13, 2017

    @bunop

    @danimag, you can't change this value with a system (system2) command, because when this function is called, a new shell environment is created, then the command is executed and finally the terminal is closed and all its variables are lost. This environment is different from Rstudio, so you can't change it, even by calling a system and put a very long process in background. The only way is by changing it permanently (but unfortunately I don't know how to change it on Mac OSX) or temporarly (using ulimit). Then all R and Rstudio session need to start from the same terminal (I found here a tip on how to launch Rstudio from a terminal on OSX, but I have no OSX so I have no idea if this trick works; if you want to try it on linux, simply type rstudio in terminal). IMHO is better to use R terminal in dedicated infrastructures when dealing with big data.

  11. danimfernandes commented on Jan 13, 2017

    @danimfernandes

    @bunop That is a good point, I did not realise that about system. Opening R or Rstudio from the terminal in linux is just like you said, so no issues there. I would gladly do that in-house, but I don't feel like seeing that as an option if I eventually publish this script. If I do maybe I could put an optional argument to whether use data.table for this or not though.

  12. bunop commented on Jan 13, 2017

    @bunop

    I understand. My interest in this discussion is because I didn't found a solution to this problem. The main point is that such problem is not related to data.table itself but on your system. I realize also that such solutions are not familiar to the basic user, and need to be avoided if possible. Take a look also at bigtable (you can find an intro here). In our case this solution doesn't work because our ped files may contain letters and this is not supported at the moment by the library. Moreover, all data need to be of the same type (like a matrix). Good luck!

  13. danimfernandes commented on Jan 13, 2017

    @danimfernandes

    Letters is all my ped files are made of, but thanks for the suggestion! :) It seems there are a few great alternatives to r-base for large datasets.

  14. costis-t commented on Jun 26, 2017

    @costis-t

    I encountered the same problem. I summarize the issue and include some additional details for future reference.

    There are two main cases:

    1. File has not a lot of columns
    2. File has indeed a lot of columns

    File has not a lot of columns

    Sometimes, there are files which have a mixed Windows and Unix line endings. I believe fread() should fix and warn about such cases.

    In my case, only the first line with the headers in the .csv had a \r\n (CRLF) in that order, the Windows standard. fread() was confused and mistakenly thought that my large file had ~300 millions of columns, since it thought that every comma-separated value was a column! To check if this is indeed the problem, enable the -A option in cat and you will see that the first line ends with a ^M:

    cat -A first_three_lines.csv 
    Timestamp,Open,High,Low,Close,Volume,Volume_(Currency),Weighted_Price^M$
    1325192180,3.237,3.237,3.237,3.237,0.3,1.6988,3.237$
    1326100360,3.1,3.1,3.1,3.1,0.623628,2.6668738,3.1$

    To solve this issue, delete the ^M characters manually with an editor like vim or sed or use the dos2unix utility (it should be available in your linux repositories),

    File has indeed a lot of columns

    In the case that the file has a lot of columns (think of millions+), in linux environments, you must increase the stack size limit of your session. This is done (read reference for potential dangers) either

    • via the shell or
    • setting them via PAM modules (recommended).

    You cannot set soft limits via the shell if they are larger than the hard limits.

    The current stack size limit is shown running ulimit -s (or ulimit -a | grep 'stack size').
    The soft limits (can temporarily be increased) cannot be larger than the hard limits.
    Setting a limit larger than the hard limit might be the cause of the bash: ulimit: stack size: cannot modify limit: Operation not permitted error when setting a very large limit.

    One other cause of this error is that, in my system, I cannot increase the limit after decreasing it:

    $ ulimit -s
    102400 # the current stack size limit
    ulimit -s 20000 # decrease to 20000 kbytes
    $ ulimit -s
    20000 # the current decreased stack size limit
    $ ulimit -s 21000
    # bash: ulimit: stack size: cannot modify limit: Operation not permitted # cannot be increased! you have to begin a new session; open a new shell

    Increase the limits with PAM

    To increase the limits permanently in your system for every new session:

    1. Add the pam_limit.so module in /etc/pam.d/login to enable it:

      cat /etc/pam.d/login
      # <truncated output>
      # session  required       /lib/security/pam_limits.so
      
    2. Increase the hard or soft stack size limit in /etc/security/limits.conf (the file itself should be well-documentated) by adding these lines:

      your_username hard stack 102400
      your_username soft stack 8192 # or more if you never want to use ulimit -s
      
    3. Reboot or reload/restart your PAM configuration/service.

    Note that even with a large limit Cstack_info() can report NA (I do not know why), without limiting fread().

     ulimit -s 102300; ulimit -s; Rscript <(echo 'Cstack_info()')
    # 102300 # The stack size limit is 102300
    # size    current  direction eval_depth 
    #   NA         NA          1          2   # stack size limit is 102300 but R reports NA
    ulimit -s 82300; ulimit -s; Rscript <(echo 'Cstack_info()')
    # 82300 # The stack size limit is 82300
    #       size    current  direction eval_depth 
    #   80061440       8320          1          2  # stack size limit is 82300 and R does not report NA

    Setting the limit appropriately, you should be able to read the file with a very large number of columns.
    In my case, setting a limit with ulimit -s 102300 and no less, I could read a file with around 13,000,000 columns!

    Errors that R/data.table may throw:

    1. When the required stack size limit is very close to the required limit by the number of the millions of columns, I got this error:

      Error: C stack usage 8036772 is too close to the limit
      
    2. When the number of columns were quite larger for the required stack size limit, I got:

      Error: segfault from C stack overflow
      
    3. R crashed when the number of columns was even larger!:

      Segmentation fault (core dumped)
      

      or

      *** caught segfault ***
      address 0x7fffce96a820, cause 'memory not mapped'
      Traceback:
      1: fread("data.csv")
      Possible actions:
      <trunccated>
      
  15. st-pasha commented on Jul 6, 2017

    @st-pasha
    Contributor

    The advice on how to increase stack size is very helpful; however latest version of fread does not attempt to allocate on the stack any more, so hopefully these tweaks should not be necessary. In particular, on my system with stack size 8192KB the example file can be read without any hiccups.

    The problem of inconsistent line endings is a separate problem, and deserves a dedicated issue.

  16. AndreV84 commented on Jan 6, 2018

    @AndreV84

    Hi, While the solution works for R. It doesn't appear to increase stack size for Rstudio-server.
    Could you provide any advise on how to get the stack increased for rstudio-serverec2 [ami instance , in particular]?

  17. bunop commented on Jan 7, 2018

    @bunop

    Hi. The point is that changing the ulimit value in a terminal affect only R instances in that terminal, as described in the discussion. This change isn't permanent, so if you relog in a different terminal or reboot the machine, you have to apply the same command. Maybe rstudio server starts from a init script, and so ulimit has no effect on it. if you have a rstudio server instance on a machine (regardless if it is a physical, virtual, o even in a docker container) the easiest solution could be change this value system wide or set this value in the same (init) script which starts Rstudio server, or start rstudio server manually in the same unlimited terminal. Hope this helps

  18. AndreV84 commented on Jan 7, 2018

    @AndreV84

    I tried to. I did set the value in a permanent way. But it wont help get Cstack_info() in Rstudio-server changed. However, it works brilliantly with R.

  19. AndreV84 commented on Jan 7, 2018

    @AndreV84

    I as well approached to stop rstudio-server and start it from /sbin/rstudio-server from the terminal. It wont help,m unfortunately.

  20. AndreV84 commented on Jan 7, 2018

    @AndreV84

    How to "set this value in the same (init) script which starts Rstudio server" ?
    I tried to execute system(ulimit -s 65535) under rstudio-server web interface neither it impacts things, as it seems to me

  21. bunop commented on Jan 7, 2018

    @bunop

    I think if you change this values system wide (and you see this value changed in each of your terminal) but rstudio server still report the same value could be that rstudio server has its ulimit configuration file or you have changed the ulimit value for a user that isn't the user which starts rstudio. If it is a issue regarding rstudio, maybe post this problem into its community may helps. You can't change this value using a system in rstudio server, I've explained why before in the discussion

  22. AndreV84 commented on Jan 7, 2018

    @AndreV84

    I tried both. I posted to community forum. Bur rstudio pro [tested with amazon ec2] wont support the Cstack increase as well, as it seems to me. I posted to Rstudio Pro forum. hopefully they will respond with some meaningful advise soon.
    References:
    https://support.rstudio.com/hc/en-us/requests/24888
    https://support.rstudio.com/hc/en-us/community/posts/115008248313-Cstack-info-value-wont-change-in-rstudio-server

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions