Skip to content

[R-Forge #4649] fread files with more than 4 billion rows #570

Description

@arunsrinivasan

Submitted by: Roby Joehanes; Assigned to: Nobody; R-Forge link

I have a huge table with 4+ billion rows. fread has an integer overflow problem and only read about 73M rows.

Activity

  1. added this to the v1.9.6 milestone on Oct 25, 2014
  2. mattdowle commented on Oct 25, 2014

    @mattdowle
    Member

    This will need data.table itself to support more than 2 billion rows.
    But the integer overflow when counting rows should be caught in the meantime.

  3. mattdowle commented on Nov 15, 2014

    @mattdowle
    Member

    The integer overflow at 2bn is now caught by e15facd.
    Pushing supporting for >2bn rows to v1.9.8

  4. modified the milestones: v1.9.8, v1.9.6 on Nov 15, 2014
  5. modified the milestones: , v1.9.8 on Mar 8, 2016
  6. removed this from the milestone on May 10, 2018
  7. robbyjo commented on Sep 15, 2026

    @robbyjo

    Proposed patch against commit a3eeb4c

    data.table-long-vector-64bit.patch
    This is made by ChatGPT Sol 5.6-Extra high under a US government license
    Extends the number of rows to R_XLEN_T_MAX = 2^52
    Edit 2: I just noticed I was the original submitter of this PR from R-forge.

  8. MichaelChirico commented on Sep 15, 2026

    @MichaelChirico
    Member

    Proposed patch against commit a3eeb4c

    data.table-long-vector-64bit.patch This is made by ChatGPT Sol 5.6-Extra high under a US government license Extends the number of rows to R_XLEN_T_MAX = 2^52 Edit 2: I just noticed I was the original submitter of this PR from R-forge.

    https://github.com/Rdatatable/data.table/blob/master/.github/CONTRIBUTING.md#important-ai-policy-

  9. robbyjo commented on Sep 15, 2026

    @robbyjo

    Proposed patch against commit a3eeb4c
    data.table-long-vector-64bit.patch This is made by ChatGPT Sol 5.6-Extra high under a US government license Extends the number of rows to R_XLEN_T_MAX = 2^52 Edit 2: I just noticed I was the original submitter of this PR from R-forge.

    https://github.com/Rdatatable/data.table/blob/master/.github/CONTRIBUTING.md#important-ai-policy-

    https://openai.com/policies/services-agreement/
    Especially point 4.1:
    4.1. Generally. Customer and Customer’s End Users may provide Input and receive Output. As between Customer and OpenAI, to the extent permitted by applicable law, Customer: (a) retains all ownership rights in Input; and (b) owns all Output. OpenAI hereby assigns to Customer all OpenAI’s right, title, and interest, if any, in and to Output.

    (Emphasis mine)

  10. jangorecki commented on Sep 15, 2026

    @jangorecki
    Member

    @robbyjo

    Regarding patch.
    Did you run our full test suite?
    did it succeeded?
    is there any runtime overhead that non-long vector are now have to carry as a result of this change?

    Regarding license and agents.
    Could you please double check with your agent that code written was trained only on MPL-2.0 compatible code? ideally let him provide you references.


    I asked my agent about that:

    https://openai.com/policies/services-agreement/
    Especially point 4.1:
    4.1. Generally. Customer and Customer’s End Users may provide Input and receive Output. As between Customer and OpenAI, to the extent permitted by applicable law, Customer: (a) retains all ownership rights in Input; and (b) owns all Output. OpenAI hereby assigns to Customer all OpenAI’s right, title, and interest, if any, in and to Output.

    does it mean if openai write code for me, I can be sure that code produced is compatible license, like GPL or MPL ? answer yes or not.

    Received: No.

    (Emphasis mine)

  11. robbyjo commented on Sep 15, 2026

    @robbyjo

    @robbyjo

    Regarding patch. Did you run our full test suite? did it succeeded? is there any runtime overhead that non-long vector are now have to carry as a result of this change?

    Regarding license and agents. Could you please double check with your agent that code written was trained only on MPL-2.0 compatible code? ideally let him provide you references.

    I asked my agent about that:

    https://openai.com/policies/services-agreement/
    Especially point 4.1:
    4.1. Generally. Customer and Customer’s End Users may provide Input and receive Output. As between Customer and OpenAI, to the extent permitted by applicable law, Customer: (a) retains all ownership rights in Input; and (b) owns all Output. OpenAI hereby assigns to Customer all OpenAI’s right, title, and interest, if any, in and to Output.

    does it mean if openai write code for me, I can be sure that code produced is compatible license, like GPL or MPL ? answer yes or not.

    Received: No.

    (Emphasis mine)

    I asked ChatGPT Astra High (this time on my own private account):
    If a code generated in ChatGPT codex, which is a patch to an MPL licensed open source code, given the services agreement (https://openai.com/policies/services-agreement/ ) and in the US government computer, can it be licensed to one compatible to MPL (to be submitted back to the MPL-licensed project)? This patch was performed as a part of official duty of US government contractors whose contract explicitly mentions that all IPs generated from the contracts belong to the US government. Please answer concisely in 2-3 sentences.

    Answer: Yes—generally, the patch can be submitted under MPL 2.0 (or the project’s required MPL-compatible contribution terms): OpenAI’s Services Agreement assigns its rights in Codex output to the customer, while MPL permits modifications and requires distributed modifications to covered files to remain under MPL. ([OpenAI]1) Because the contractor agreement expressly assigns work-product IP to the U.S. Government, the Government would ordinarily be the relevant rights holder able to authorize that contribution, subject to the agency’s contract, security, and release policies; Codex output can also carry third-party/open-source licensing obligations that should be checked before submission. ([openai.com]2)

    Source: https://chatgpt.com/share/6aa950b5-b75c-83ea-ac09-7dc8134ef77d

    About the patch: I ran the full test suite through ChatGPT and it succeeded. I can rerun the suite manually (if you accept the licensing answer; if not, I'm using my own privately and it works)
    Overhead: Have not check. Will do. (Subject to an agreeable licensing answer)

  12. aitap commented on Sep 15, 2026

    @aitap
    Member

    Dear Roby Joehanes,

    Thank you for extending the effort to put a patch together and offer us the code! Currently, the patch has other problems beyond the copyright concerns («Codex output can also carry third-party/open-source licensing obligations that should be checked»), so we cannot accept it. We'll develop a comprehensive solution that will cover all bases that we care about; meanwhile, your fork works well for you. I recommend talking to Håkon Tjeldnes to avoid duplication of effort.

  13. robbyjo commented on Sep 15, 2026

    @robbyjo

    Dear Roby Joehanes,

    Thank you for extending the effort to put a patch together and offer us the code! Currently, the patch has other problems beyond the copyright concerns («Codex output can also carry third-party/open-source licensing obligations that should be checked»), so we cannot accept it. We'll develop a comprehensive solution that will cover all bases; meanwhile, your fork works well for you. I recommend talking to Håkon Tjeldnes to avoid duplication of effort.

    I did follow up discussion that absolves me from any copyright entanglement: https://chatgpt.com/share/6aa96f6f-00fc-83ea-92e7-42d1fb2b1c96


    I see no additional GPL obligation arising from R source code in this patch. The patch appears to be an independently implemented extension of data.table's existing MPL-2.0 code using R's documented/public long-vector and ALTREP APIs, rather than a port of R's GPL implementation.

    I would therefore be comfortable describing the output-level provenance as: no identifiable third-party implementation copied from R; R-derived material is limited to public API usage, documented semantics, and generic programming idioms. What cannot truthfully be certified is the maintainer's request that the model was “trained only on MPL-2.0 compatible code”—model training provenance is a separate question and cannot be established from the patch.


    Edit: Even Håkon Tjeldnes used Astra xhigh. If you won't accept any AI-made code, then you could not accept his as well.

  14. jangorecki commented on Sep 15, 2026

    @jangorecki
    Member

    Just for the record, I am not against using AI.
    Asking AI to identify which parts of code need to be updated when I change signature of a function is different than making AI write the change.

    Our current AI policy could be changed, but that would require data.table team to look into it. Especially Matt Dowle's take on it, as he offered this project to the community.
    Allowing AI contribution results in more maintainer time required.
    Whoever needs to review or merge PR also needs to understand it, unlike the person who is submitting it, who can produce big and complex patches in minutes.

  15. Roleren commented on Sep 15, 2026

    @Roleren

    I am a bit with aitap on this issue, even though the new astra model is amazing, and I manually read the tests it makes, I think all of us need a few months to digest what the hell to do with all this magic we can now do, to make some rigorous guidelines etc, the top models can now do more work in 15 minutes than I can do in a month of spare time coding (especially the C code). There is so much potential here for us package developers to improve so much for the whole community, but full human verification is basically impossible.

    I guess the way to go would be something similar to the Copilot code review (to get a "verified stamp") combined with some minimum commits / owned repos from the PR author with filters like (not accounts that are < 6 months old etc) to make it more credible.

    But for now, we can just work on our forks @robbyjo (+ whoever else want to help us) and when we all feel a bit more comfortable with ai written code we can discuss a possible pull request to the main repo.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions