Skip to content

(feat) Additional Arrow support #1785

Description

@Fred-Wu

Additional Arrow support

With the current on demand data loading, I think Data Viewer could extend existing Arrow table support to lazy Arrow-backed data sources:

  • Dataset and its subclasses, including FileSystemDataset
  • arrow_dplyr_query

This allows file-backed Arrow datasets and dplyr queries on Arrow data to be opened directly with View() without first collecting the full dataset into memory. Rows are fetched on demand as they are requested by the Data Viewer, allowing large Arrow datasets to be browsed efficiently.

This should only be a small change in the data viewer model.

Activity

  1. added this to the 3.x milestone on Sep 27, 2026
  2. added theissue type on Sep 27, 2026
  3. grantmcdermott commented on Sep 27, 2026

    @grantmcdermott
    Contributor

    Could we not implement via a more lightweight solution like nanoarrow? https://arrow.apache.org/nanoarrow/latest/index.html

  4. Fred-Wu commented on Sep 28, 2026

    @Fred-Wu
    ContributorAuthor

    Could we not implement via a more lightweight solution like nanoarrow? https://arrow.apache.org/nanoarrow/latest/index.html

    It looks like nanoarrow is mainly designed for data streaming. It still relies on arrow for Dataset and arrow_dplyr_query. For our data viewer, we also need efficient forward and backward scrolling, filtering, sorting, and caching, so arrow seems better fitted here. Type conversion might also be an issue with nanoarrow.

    For those files sit on disc rather than in memory, I think some trade-offs are expected. For example, loading data might be slower compared to in-memory data, especially after sorting.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions