Skip to content

csv_to_disk.frame writing error on relative path #340

Description

@matthewgson

I think I found a bug case where csv_to_disk.frame function does not recognize relative outdir. The solution was to provide absolute path on outdir.

Here's minimal replicable example.

setwd('~/Documents/')
fwrite(iris, 'iris.csv')
iris.df = csv_to_disk.frame('iris.csv',
                            'iris.df',
                            in_chunk_size = 10,
                            select = c('Sepal.Length','Petal.Length','Species') # fread arg
)


 ----------------------------------------------------- 
Stage 1 of 2: splitting the file iris.csv into smallers files:
Destination: /var/folders/lc/8s45bh7s7mz06tb7k7fp74p40000gn/T//RtmpRoVeur/file81b042b828b6
 ----------------------------------------------------- 
Stage 1 of 2 took: 0.004s elapsed (0.001s cpu)
 ----------------------------------------------------- 
Stage 2 of 2: Converting the smaller files into disk.frame
 ----------------------------------------------------- 
csv_to_disk.frame: Reading multiple input files.
Please use `colClasses = `  to set column types to minimize the chance of a failed read
=================================================

 ----------------------------------------------------- 
-- Converting CSVs to disk.frame -- Stage 1 of 2:

Converting 16 CSVs to 4 disk.frames each consisting of 4 chunks

 Progress: ─────────────────────────────────────────────────────────────────────────────────────────────────────────── 100%-- Converting CSVs to disk.frame -- Stage 1 or 2 took: 1.014s elapsed (0.114s cpu)
 ----------------------------------------------------- 
 
 ----------------------------------------------------- 
-- Converting CSVs to disk.frame -- Stage 2 of 2:

Row-binding the 4 disk.frames together to form one large disk.frame:
Creating the disk.frame at iris.df

Appending disk.frames: 
Error in fst::write_fst(purrr::map_dfr(full_paths1, ~fst::read_fst(.x)),  : 
  There was an error creating the file, please check path

sessionInfo()
R version 4.0.3 (2020-10-10)
Platform: x86_64-apple-darwin17.0 (64-bit)
Running under: macOS 11.4

packageVersion('disk.frame')
[1] ‘0.5.0’

I experienced same error on my windows machine too.

Activity

  1. xiaodaigh commented on Jun 6, 2021

    @xiaodaigh
    Collaborator
    bug.mp4

    I can't replicate the error see video.

  2. matthewgson commented on Jun 6, 2021

    @matthewgson
    Author

    Thanks for prompt feedback!
    It's my bad that I did not run above minimum example on windows machine. I confirm that above example runs fine on windows. Only from my Mac I can replicate above result.

    The original code that I was trying to run on windows gave me the error message so I assumed above example would do on windows. Differences I can point out are merely 1) writing on Dropbox-linked folder, 2) using inmapfn though. I'll update if I can make replicable example on windows.

  3. xiaodaigh commented on Jun 6, 2021

    @xiaodaigh
    Collaborator

    there's no platform-specific code IIRC. So might be something else you are experiencing.. thanks for the MWE

  4. matthewgson commented on Jun 6, 2021

    @matthewgson
    Author

    It's a bit strange, but here's the complete log I have from windows.
    On the first run it generated error, but on the second run it ran smoothly... I restarted R session before the second run.

    R version 4.0.2 (2020-06-22) -- "Taking Off Again"
    Copyright (C) 2020 The R Foundation for Statistical Computing
    Platform: x86_64-w64-mingw32/x64 (64-bit)
    
    R is free software and comes with ABSOLUTELY NO WARRANTY.
    You are welcome to redistribute it under certain conditions.
    Type 'license()' or 'licence()' for distribution details.
    
    R is a collaborative project with many contributors.
    Type 'contributors()' for more information and
    'citation()' on how to cite R or R packages in publications.
    
    Type 'demo()' for some demos, 'help()' for on-line help, or
    'help.start()' for an HTML browser interface to help.
    Type 'q()' to quit R.
    
    > library(disk.frame)
    Loading required package: dplyr
    
    Attaching package: ‘dplyr’
    
    The following objects are masked from ‘package:stats’:
    
        filter, lag
    
    The following objects are masked from ‘package:base’:
    
        intersect, setdiff, setequal, union
    
    Loading required package: purrr
    Registered S3 method overwritten by 'pryr':
      method      from
      print.bytes Rcpp
    
    ## Message from disk.frame:
    We have 1 workers to use with disk.frame.
    To change that, use setup_disk.frame(workers = n) or just setup_disk.frame() to use the defaults.
    
    
    It is recommended that you run the following immediately to set up disk.frame with multiple workers in order to parallelize your operations:
    
    
    # this will set up disk.frame with multiple workers
    setup_disk.frame()
    # this will allow unlimited amount of data to be passed from worker to worker
    options(future.globals.maxSize = Inf)
    
    
    
    
    
    Attaching package: ‘disk.frame’
    
    The following objects are masked from ‘package:purrr’:
    
        imap, imap_dfr, map, map2
    
    The following objects are masked from ‘package:base’:
    
        colnames, ncol, nrow
    
    Warning messages:
    1: package ‘disk.frame’ was built under R version 4.0.4 
    2: package ‘dplyr’ was built under R version 4.0.5 
    > setup_disk.frame(workers = 8)
    The number of workers available for disk.frame is 8
    > options(future.globals.maxSize = Inf,  # Unlimited data communication
    +         future.rng.onMisuse = 'ignore') # Ignore random seeds
    > setwd('C:/Users/gunsu.son/Documents/')
    > dir.create('iris_folder')
    
    > dir.create('iris_folder')
    Warning message:
    In dir.create("iris_folder") : 'iris_folder' already exists
    > data.table::fwrite(iris, 'iris_folder/iris.csv')
    > iris.df = csv_to_disk.frame('iris_folder/iris.csv',
    +                             'iris.df',
    +                             inmapfn = function(chunk){
    +                               chunk[, Newvar := 1]
    +                             },
    +                             in_chunk_size = 10,
    +                             compress=100,
    +                             select = c('Sepal.Length','Petal.Length','Species'))
     ----------------------------------------------------- 
    Stage 1 of 2: splitting the file iris_folder/iris.csv into smallers files:
    Destination: C:\Users\gunsu.son\AppData\Local\Temp\RtmpkJ5jMh\file2f2034f45e42
     ----------------------------------------------------- 
    Stage 1 of 2 took: 0.000s elapsed (0.000s cpu)
     ----------------------------------------------------- 
    Stage 2 of 2: Converting the smaller files into disk.frame
     ----------------------------------------------------- 
    csv_to_disk.frame: Reading multiple input files.
    Please use `colClasses = `  to set column types to minimize the chance of a failed read
    =================================================
    
     ----------------------------------------------------- 
    -- Converting CSVs to disk.frame -- Stage 1 of 2:
    
    Converting 16 CSVs to 4 disk.frames each consisting of 4 chunks
    
    -- Converting CSVs to disk.frame -- Stage 1 or 2 took: 11.0s elapsed (0.160s cpu)
     ----------------------------------------------------- 
     
     ----------------------------------------------------- 
    -- Converting CSVs to disk.frame -- Stage 2 of 2:
    
    Row-binding the 4 disk.frames together to form one large disk.frame:
    Creating the disk.frame at iris.df
    
    Appending disk.frames: 
    Error in fst::write_fst(purrr::map_dfr(full_paths1, ~fst::read_fst(.x)),  : 
      There was an error creating the file, please check path
    
    
    Restarting R session...
    
    > library(disk.frame)
    Loading required package: dplyr
    
    Attaching package: ‘dplyr’
    
    The following objects are masked from ‘package:stats’:
    
        filter, lag
    
    The following objects are masked from ‘package:base’:
    
        intersect, setdiff, setequal, union
    
    Loading required package: purrr
    Registered S3 method overwritten by 'pryr':
      method      from
      print.bytes Rcpp
    
    ## Message from disk.frame:
    We have 1 workers to use with disk.frame.
    To change that, use setup_disk.frame(workers = n) or just setup_disk.frame() to use the defaults.
    
    
    It is recommended that you run the following immediately to set up disk.frame with multiple workers in order to parallelize your operations:
    
    
    
    # this will set up disk.frame with multiple workers
    setup_disk.frame()
    # this will allow unlimited amount of data to be passed from worker to worker
    options(future.globals.maxSize = Inf)
    
    
    
    
    
    Attaching package: ‘disk.frame’
    
    The following objects are masked from ‘package:purrr’:
    
        imap, imap_dfr, map, map2
    
    The following objects are masked from ‘package:base’:
    
        colnames, ncol, nrow
    
    Warning messages:
    1: package ‘disk.frame’ was built under R version 4.0.4 
    2: package ‘dplyr’ was built under R version 4.0.5 
    > setup_disk.frame(workers = 8)
    The number of workers available for disk.frame is 8
    > options(future.globals.maxSize = Inf,  # Unlimited data communication
    +         future.rng.onMisuse = 'ignore') # Ignore random seeds warnings
    > 
    > setwd('C:/Users/gunsu.son/Documents/')
    > dir.create('iris_folder')
    Warning message:
    In dir.create("iris_folder") : 'iris_folder' already exists
    > data.table::fwrite(iris, 'iris_folder/iris.csv')
    > 
    > iris.df = csv_to_disk.frame('iris_folder/iris.csv',
    +                             'iris.df',
    +                             in_chunk_size = 10,
    +                             compress=100,
    +                             select = c('Sepal.Length','Petal.Length','Species'))
     ----------------------------------------------------- 
    Stage 1 of 2: splitting the file iris_folder/iris.csv into smallers files:
    Destination: C:\Users\gunsu.son\AppData\Local\Temp\Rtmp0IZENN\file251c1d296ffb
     ----------------------------------------------------- 
    Stage 1 of 2 took: 0.000s elapsed (0.000s cpu)
     ----------------------------------------------------- 
    Stage 2 of 2: Converting the smaller files into disk.frame
     ----------------------------------------------------- 
    csv_to_disk.frame: Reading multiple input files.
    Please use `colClasses = `  to set column types to minimize the chance of a failed read
    =================================================
    
     ----------------------------------------------------- 
    -- Converting CSVs to disk.frame -- Stage 1 of 2:
    
    Converting 16 CSVs to 4 disk.frames each consisting of 4 chunks
    
    -- Converting CSVs to disk.frame -- Stage 1 or 2 took: 11.0s elapsed (0.110s cpu)
     ----------------------------------------------------- 
     
     ----------------------------------------------------- 
    -- Converting CSVs to disk.frame -- Stage 2 of 2:
    
    Row-binding the 4 disk.frames together to form one large disk.frame:
    Creating the disk.frame at iris.df
    
    Appending disk.frames: 
    Stage 2 of 2 took: 0.080s elapsed (0.050s cpu)
     ----------------------------------------------------- 
    Stage 1 & 2 in total took: 11.1s elapsed (0.170s cpu)
    Stage 2 of 2 took: 11.3s elapsed (0.180s cpu)
     ----------------------------------------------------- 
    Stage 2 & 2 took: 11.3s elapsed (0.180s cpu)
     -----------------------------------------------------
    
    
  5. matthewgson commented on Jun 6, 2021

    @matthewgson
    Author

    Ok.... I'm a bit confused why this is occurring, but I think this code would replicate the issue I was encountering on windows.

    Terminate R, fresh new R session

    library(disk.frame)
    setup_disk.frame(workers = 8)
    options(future.globals.maxSize = Inf,  # Unlimited data communication
            future.rng.onMisuse = 'ignore') # Ignore random seeds warnings
    
    setwd('C:/Users/gunsu.son/Documents/')
    dir.create('iris_folder')
    data.table::fwrite(iris, 'iris_folder/iris.csv')
    
    iris.df = csv_to_disk.frame('iris_folder/iris.csv',
                                'iris.df',
                                in_chunk_size = 10,
                                compress=100,
                                select = c('Sepal.Length','Petal.Length','Species'))
    

    When above code was ran the second time, it worked well. However, when I terminate R and run the code again, it produced error again.

  6. xiaodaigh commented on Jun 6, 2021

    @xiaodaigh
    Collaborator

    not sure what's happening. there are sometimes weird issues with folder access. try not to write to dropbox directly and see if that helps? I suspect it's not a disk.frame issue but software or OS issue. I will close it for now unless you have more issues with it.

  7. matthewgson commented on Jun 6, 2021

    @matthewgson
    Author

    @xiaodaigh Could you confirm if above example produces the issue on your side?

    Edit: my hunch is that it's related to setup_disk.frame(workers = 8) - future workers are not fully initiated on the first run of the code, but it seems to work on the second time it was ran..

  8. xiaodaigh commented on Jun 6, 2021

    @xiaodaigh
    Collaborator

    It ran successfully for me. Can you try a different folder?

  9. matthewgson commented on Jun 6, 2021

    @matthewgson
    Author

    I guess the issue belongs to my settings then, not disk.frame issue. It’s still strange for me to see it happening sometimes.

  10. matthewgson commented on Jun 6, 2021

    @matthewgson
    Author

    @xiaodaigh
    Well... I think I figured out why this happened. Seems it's related to the working directory of future workers. When setup_disk.frame(workers=8) was called, it produces R sessions based on the current working directory of master R session, and their working directory is not updated when setwd was called later. So they are confused when they are asked to write fst on outdir in relative path. Calling setup_disk.frame once again updates the cwd and that resolves the issue.

    I think above example worked for you because your home directory is set at 'C:/Users/USERNAME/Documents/' while my machine has somewhere else when it first freshly launched.

    Wish I had this in my mind in the first place!

  11. changed the title [-]csv_to_disk.frame writing error on relative path when fread arg was used[/-] [+]csv_to_disk.frame writing error on relative path[/+] on Jun 6, 2021
  12. xiaodaigh commented on Jun 7, 2021

    @xiaodaigh
    Collaborator

    Glad it's solved. I suspected something like this but I had forgotten hat I had run setup_disk.frame again after doing setwd. In general, setwd is messy and I never do ti.

  13. matthewgson commented on Jun 7, 2021

    @matthewgson
    Author

    @xiaodaigh Would it be possible to add a helpful error message? It will help users to grasp best practices when using disk.frame, I believe. And tremendous thank you for creating this awesome package, it's a lifesaver for me.

  14. xiaodaigh commented on Jun 7, 2021

    @xiaodaigh
    Collaborator

    yeah see #341

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions