Skip to content

Add CSVY support for fread() #1701

Description

@arunsrinivasan

The rio already allows reading of csvy formats by relying on the helper package (same author) csvy. It'd be great to leverage its features to read/write.

Activity

  1. MichaelChirico commented on Mar 3, 2018

    @MichaelChirico
    Member

    Just documenting how the .csvy reader of rio works:

    1. rio encounters a .csvy file and dispatches to .import.rio_csvy
    2. .import.rio_csvy calls read_csvy from the csvy package.
    3. read_csvy uses readLines to digest the file, then uses grep to identify the YAML header.
    4. yaml.load from the yaml package is called to convert the YAML content string to a list of component parts. This is already implemented in C so is presumably efficient.
    5. read_csv then applies paste(., collapse = '\n') to the non-YAML portion of the file and can read it with fread.
    6. Content of YAML header is applied to the output.

    Major inefficiencies are:

    • Using readLines to digest the whole file (slow)
    • pasteing the file back into a format for fread to tackle only after stripping out the YAML part
    • Some parts of the YAML header are intended to assist fread, but the YAML data is not fed into fread.

    My proposed solution (if we decide to tackle this):

    1. Stream lines of the file until the end of the YAML header is reached (or until an out-of-format line is reached -> stop). I'm not sure the most efficient way to pass files line-by-line in R, or if we'll have to implement that ourselves in C. Keep track of the # of lines fed. I see this from StackOverflow: https://stackoverflow.com/q/9871307/3576984
    2. Add yaml package to Suggests and rely on that to parse the YAML info
    3. Extract any info relevant to fread itself (especially/most importantly colClasses; I'll have to read the csvy standard to see how open-ended the rest is)
    4. fread the remainder of the file, using skip to jump past the YAML section, and including relevant info from the YAML.

    Remaining API Q for me are: (1) do we try and detect YAML automatically (harder & beyond me to implement currently), or simply add a yaml (or similarly named) argument to fread and rely on user input? (2) What of fwrite?

    For (1) I lean towards the latter primarily out of laziness (more sophisticated -- doesn't seem to pass the cost/benefit test given limited user requests for this feature). We can revisit in a future issue if this becomes more popular, I suppose. No opinions on (2), I only include it since we seem to be aiming to keep fread and fwrite as each others' inverse functions.

  2. HughParsonage commented on Mar 3, 2018

    @HughParsonage
    Member

    I'm not sure the most efficient way to pass files line-by-line in R

    fread(file, sep = NULL) (in dev) ?

    Also readLines is likely to be much faster in R 3.5.0, possibly as fast as fread or readr::read_lines.

  3. MichaelChirico commented on Mar 3, 2018

    @MichaelChirico
    Member

    @HughParsonage that still leaves the issue of reading the file twice. The idea of streaming lines is to examine the file line-by-line and only read in, say, 10-20 lines of YAML metadata using readLines (or fread, or whatever) before parsing that and deploying fread on (presumably) the bulk of the file which follows the header

    (anyway good to know they're finally getting around to improving readLines, it's silly how slow it is considering how minimal its responsibilities are)

  4. MichaelChirico commented on Mar 3, 2018

    @MichaelChirico
    Member
  5. jangorecki commented on Mar 5, 2018

    @jangorecki
    Member

    I would avoid extra dependency, even in suggests, and use some helper function to extract fields from yaml header. Similar way as we would process DESCRIPTION file. Fwrite should be able to write csvy header the way that fread can read it.

  6. MichaelChirico commented on Mar 5, 2018

    @MichaelChirico
    Member

    as yaml can be arbitrarily nested I didn't see a need to reinvent the wheel.

    especially as it seems at the moment to be a rather limited use case -- happy to revisit if this format takes off.

  7. added this to the 1.12.0 milestone on Jun 26, 2018
  8. modified the milestones: 1.12.0, 1.12.2 on Jan 5, 2019
  9. removed this from the 1.12.2 milestone on Jan 14, 2019
  10. added a commit that references this issue on Feb 10, 2019
    acf13b6
  11. added this to the 1.12.4 milestone on May 2, 2019
  12. MichaelChirico commented on May 3, 2019

    @MichaelChirico
    Member

    fwrite support not done yet... will file separately

  13. changed the title [-]Add CSVY support for fread() and fwrite()[/-] [+]Add CSVY support for fread()[/+] on May 3, 2019
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Type

No type

Projects

No projects

    Milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions