Repository navigation
csv_to_disk.frame writing error on relative path #340
Description
Activity
bug.mp4
I can't replicate the error see video.
Thanks for prompt feedback!
It's my bad that I did not run above minimum example on windows machine. I confirm that above example runs fine on windows. Only from my Mac I can replicate above result.The original code that I was trying to run on windows gave me the error message so I assumed above example would do on windows. Differences I can point out are merely 1) writing on Dropbox-linked folder, 2) using inmapfn though. I'll update if I can make replicable example on windows.
there's no platform-specific code IIRC. So might be something else you are experiencing.. thanks for the MWE
It's a bit strange, but here's the complete log I have from windows.
On the first run it generated error, but on the second run it ran smoothly... I restarted R session before the second run.R version 4.0.2 (2020-06-22) -- "Taking Off Again" Copyright (C) 2020 The R Foundation for Statistical Computing Platform: x86_64-w64-mingw32/x64 (64-bit) R is free software and comes with ABSOLUTELY NO WARRANTY. You are welcome to redistribute it under certain conditions. Type 'license()' or 'licence()' for distribution details. R is a collaborative project with many contributors. Type 'contributors()' for more information and 'citation()' on how to cite R or R packages in publications. Type 'demo()' for some demos, 'help()' for on-line help, or 'help.start()' for an HTML browser interface to help. Type 'q()' to quit R. > library(disk.frame) Loading required package: dplyr Attaching package: ‘dplyr’ The following objects are masked from ‘package:stats’: filter, lag The following objects are masked from ‘package:base’: intersect, setdiff, setequal, union Loading required package: purrr Registered S3 method overwritten by 'pryr': method from print.bytes Rcpp ## Message from disk.frame: We have 1 workers to use with disk.frame. To change that, use setup_disk.frame(workers = n) or just setup_disk.frame() to use the defaults. It is recommended that you run the following immediately to set up disk.frame with multiple workers in order to parallelize your operations: # this will set up disk.frame with multiple workers setup_disk.frame() # this will allow unlimited amount of data to be passed from worker to worker options(future.globals.maxSize = Inf) Attaching package: ‘disk.frame’ The following objects are masked from ‘package:purrr’: imap, imap_dfr, map, map2 The following objects are masked from ‘package:base’: colnames, ncol, nrow Warning messages: 1: package ‘disk.frame’ was built under R version 4.0.4 2: package ‘dplyr’ was built under R version 4.0.5 > setup_disk.frame(workers = 8) The number of workers available for disk.frame is 8 > options(future.globals.maxSize = Inf, # Unlimited data communication + future.rng.onMisuse = 'ignore') # Ignore random seeds > setwd('C:/Users/gunsu.son/Documents/') > dir.create('iris_folder') > dir.create('iris_folder') Warning message: In dir.create("iris_folder") : 'iris_folder' already exists > data.table::fwrite(iris, 'iris_folder/iris.csv') > iris.df = csv_to_disk.frame('iris_folder/iris.csv', + 'iris.df', + inmapfn = function(chunk){ + chunk[, Newvar := 1] + }, + in_chunk_size = 10, + compress=100, + select = c('Sepal.Length','Petal.Length','Species')) ----------------------------------------------------- Stage 1 of 2: splitting the file iris_folder/iris.csv into smallers files: Destination: C:\Users\gunsu.son\AppData\Local\Temp\RtmpkJ5jMh\file2f2034f45e42 ----------------------------------------------------- Stage 1 of 2 took: 0.000s elapsed (0.000s cpu) ----------------------------------------------------- Stage 2 of 2: Converting the smaller files into disk.frame ----------------------------------------------------- csv_to_disk.frame: Reading multiple input files. Please use `colClasses = ` to set column types to minimize the chance of a failed read ================================================= ----------------------------------------------------- -- Converting CSVs to disk.frame -- Stage 1 of 2: Converting 16 CSVs to 4 disk.frames each consisting of 4 chunks -- Converting CSVs to disk.frame -- Stage 1 or 2 took: 11.0s elapsed (0.160s cpu) ----------------------------------------------------- ----------------------------------------------------- -- Converting CSVs to disk.frame -- Stage 2 of 2: Row-binding the 4 disk.frames together to form one large disk.frame: Creating the disk.frame at iris.df Appending disk.frames: Error in fst::write_fst(purrr::map_dfr(full_paths1, ~fst::read_fst(.x)), : There was an error creating the file, please check pathRestarting R session... > library(disk.frame) Loading required package: dplyr Attaching package: ‘dplyr’ The following objects are masked from ‘package:stats’: filter, lag The following objects are masked from ‘package:base’: intersect, setdiff, setequal, union Loading required package: purrr Registered S3 method overwritten by 'pryr': method from print.bytes Rcpp ## Message from disk.frame: We have 1 workers to use with disk.frame. To change that, use setup_disk.frame(workers = n) or just setup_disk.frame() to use the defaults. It is recommended that you run the following immediately to set up disk.frame with multiple workers in order to parallelize your operations: # this will set up disk.frame with multiple workers setup_disk.frame() # this will allow unlimited amount of data to be passed from worker to worker options(future.globals.maxSize = Inf) Attaching package: ‘disk.frame’ The following objects are masked from ‘package:purrr’: imap, imap_dfr, map, map2 The following objects are masked from ‘package:base’: colnames, ncol, nrow Warning messages: 1: package ‘disk.frame’ was built under R version 4.0.4 2: package ‘dplyr’ was built under R version 4.0.5 > setup_disk.frame(workers = 8) The number of workers available for disk.frame is 8 > options(future.globals.maxSize = Inf, # Unlimited data communication + future.rng.onMisuse = 'ignore') # Ignore random seeds warnings > > setwd('C:/Users/gunsu.son/Documents/') > dir.create('iris_folder') Warning message: In dir.create("iris_folder") : 'iris_folder' already exists > data.table::fwrite(iris, 'iris_folder/iris.csv') > > iris.df = csv_to_disk.frame('iris_folder/iris.csv', + 'iris.df', + in_chunk_size = 10, + compress=100, + select = c('Sepal.Length','Petal.Length','Species')) ----------------------------------------------------- Stage 1 of 2: splitting the file iris_folder/iris.csv into smallers files: Destination: C:\Users\gunsu.son\AppData\Local\Temp\Rtmp0IZENN\file251c1d296ffb ----------------------------------------------------- Stage 1 of 2 took: 0.000s elapsed (0.000s cpu) ----------------------------------------------------- Stage 2 of 2: Converting the smaller files into disk.frame ----------------------------------------------------- csv_to_disk.frame: Reading multiple input files. Please use `colClasses = ` to set column types to minimize the chance of a failed read ================================================= ----------------------------------------------------- -- Converting CSVs to disk.frame -- Stage 1 of 2: Converting 16 CSVs to 4 disk.frames each consisting of 4 chunks -- Converting CSVs to disk.frame -- Stage 1 or 2 took: 11.0s elapsed (0.110s cpu) ----------------------------------------------------- ----------------------------------------------------- -- Converting CSVs to disk.frame -- Stage 2 of 2: Row-binding the 4 disk.frames together to form one large disk.frame: Creating the disk.frame at iris.df Appending disk.frames: Stage 2 of 2 took: 0.080s elapsed (0.050s cpu) ----------------------------------------------------- Stage 1 & 2 in total took: 11.1s elapsed (0.170s cpu) Stage 2 of 2 took: 11.3s elapsed (0.180s cpu) ----------------------------------------------------- Stage 2 & 2 took: 11.3s elapsed (0.180s cpu) -----------------------------------------------------Ok.... I'm a bit confused why this is occurring, but I think this code would replicate the issue I was encountering on windows.
Terminate R, fresh new R session
library(disk.frame) setup_disk.frame(workers = 8) options(future.globals.maxSize = Inf, # Unlimited data communication future.rng.onMisuse = 'ignore') # Ignore random seeds warnings setwd('C:/Users/gunsu.son/Documents/') dir.create('iris_folder') data.table::fwrite(iris, 'iris_folder/iris.csv') iris.df = csv_to_disk.frame('iris_folder/iris.csv', 'iris.df', in_chunk_size = 10, compress=100, select = c('Sepal.Length','Petal.Length','Species'))When above code was ran the second time, it worked well. However, when I terminate R and run the code again, it produced error again.
not sure what's happening. there are sometimes weird issues with folder access. try not to write to dropbox directly and see if that helps? I suspect it's not a disk.frame issue but software or OS issue. I will close it for now unless you have more issues with it.
@xiaodaigh Could you confirm if above example produces the issue on your side?
Edit: my hunch is that it's related to
setup_disk.frame(workers = 8)- future workers are not fully initiated on the first run of the code, but it seems to work on the second time it was ran..It ran successfully for me. Can you try a different folder?
I guess the issue belongs to my settings then, not disk.frame issue. It’s still strange for me to see it happening sometimes.
@xiaodaigh
Well... I think I figured out why this happened. Seems it's related to the working directory offutureworkers. Whensetup_disk.frame(workers=8)was called, it produces R sessions based on the current working directory of master R session, and their working directory is not updated whensetwdwas called later. So they are confused when they are asked to write fst onoutdirin relative path. Callingsetup_disk.frameonce again updates the cwd and that resolves the issue.I think above example worked for you because your home directory is set at
'C:/Users/USERNAME/Documents/'while my machine has somewhere else when it first freshly launched.Wish I had this in my mind in the first place!
- changed the title
[-]csv_to_disk.frame writing error on relative path when fread arg was used[/-][+]csv_to_disk.frame writing error on relative path[/+]on Jun 6, 2021 Glad it's solved. I suspected something like this but I had forgotten hat I had run
setup_disk.frameagain after doingsetwd. In general,setwdis messy and I never do ti.@xiaodaigh Would it be possible to add a helpful error message? It will help users to grasp best practices when using
disk.frame, I believe. And tremendous thank you for creating this awesome package, it's a lifesaver for me.yeah see #341
Reacted by Matthew Son
I think I found a bug case where
csv_to_disk.framefunction does not recognize relativeoutdir. The solution was to provide absolute path onoutdir.Here's minimal replicable example.
I experienced same error on my windows machine too.