Skip to content

reading multiple csv files #323

Description

@maskegger

Hi, Thanks for the work on this package!

I was wondering if there's a way to get filenames as a column in disk.frame. I have multiple csv files which need to be pre-filtered before importing. I'm using inmapfn option in the csv_to_disk.frame() for this. Is there also any way I can get the filename as a column in diskframe (possibly within inmapfn)? I have looked at other issues and documentation but could not find this. I can't share the data unfortunately so I have copied my function code below. Thanks!

css_rbind_df <- function(files, outdir) {
  
  css_india_raw_df <- csv_to_disk.frame(files,
                                        outdir = outdir,
                                        colClasses = list(character = c("GID_1", "NAME_1", "region_agg")),
                                        inmapfn = function(chunk) {
                                          chunk[country_agg == "USA"]
                                          }
                                    )
  
  return(css_india_raw_df)
}

Activity

  1. xiaodaigh commented on Jan 12, 2021

    @xiaodaigh
    Collaborator

    Currently this is not possible EASILY. I thought about adding this feature a while ago and probably should.

    list_of_disk.frames lapply(files, function(a_file) {
        css_india_raw_df <- csv_to_disk.frame(a_file,
                                            outdir = outdir,
                                            colClasses = list(character = c("GID_1", "NAME_1", "region_agg")),
                                            inmapfn = function(chunk) {
                                              chunk[country_agg == "USA", filename := a_file]
                                              }
                                        )
        }  
      return(css_india_raw_df)
    }
    
    fnl_disk.frame = rbindlist.disk.frame(list_of_disk.frames, outdir = outdir)
    

    Something like this might work.

  2. maskegger commented on Jan 12, 2021

    @maskegger
    Author

    Thanks ZJ. I ended up writing a function something along this lines but doesn't this hamper performance and is it a good approach if you have large number of daily log files to import (mostly in similar format but not necessarily)? It would be good to have this integrated within csv_to_disk.frame() like you mentioned.

    Separately, it would be useful if you could add more documentation (a vignette perhaps?) on disk.frame's integration with drake. There's some info in the drake manual on this but maybe not enough. I understand I may be asking too much from you!

  3. xiaodaigh commented on Jan 13, 2021

    @xiaodaigh
    Collaborator

    (a vignette perhaps?) on disk.frame's integration with drake. There's some info in the drake manual on this but maybe not enough. See this issue #324

    Yeah.

    good approach if you have large number of daily log files to import (mostly in similar format but not necessarily)? It would be good to have this integrated within csv_to_disk.frame() like you mentioned.

    I think i will add this feature. seems like a good feature.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions