Skip to content

select via pattern in fread #2066

Description

@MichaelChirico

I'm trying to read some files where the column names (for the most part) are followed by two digits indicating the year. I'd like to select only a few of the >100 columns from each file, but currently have to do a lot more leg work to get this to work because though the column pattern is the same across files, the full name is not.

Specifically, I'm reading some csv-ified files from here.

I might want the SCHNAM (school name) column for several years; they would be stored as, e.g., SCHNAM06 for 2006-07, SCHNAM07 for 2007-08, SCHNAM08 for 2008-09, and so on.

It would be great to do fread(file_name, select = grep('SCHNAM', .)). But instead I have to do something along the lines of readLines(file_name, n = 1L) %>% strsplit(',') %>% el %>% grep('SCHNAM', .) first.

Activity

  1. lmullany commented on Jan 27, 2023

    @lmullany

    agree, this would be most helpful. This is my current work around:

    find_cols <- function(pth,re,...) {
      cols = names(fread(pth, nrows=0))
      cols[grepl(re,cols,...)]
    }
    pth = "dataset_with_hundreds_of_columns.csv"
    df = fread(pth, select=find_cols(pth,"<my_awesome_regex>")
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions