Repository navigation
[R-Forge #4649] fread files with more than 4 billion rows #570
Description
Activity
This will need data.table itself to support more than 2 billion rows.
But the integer overflow when counting rows should be caught in the meantime.The integer overflow at 2bn is now caught by e15facd.
Pushing supporting for >2bn rows to v1.9.8Proposed patch against commit a3eeb4c
data.table-long-vector-64bit.patch
This is made by ChatGPT Sol 5.6-Extra high under a US government license
Extends the number of rows to R_XLEN_T_MAX = 2^52
Edit 2: I just noticed I was the original submitter of this PR from R-forge.Proposed patch against commit a3eeb4c
data.table-long-vector-64bit.patch This is made by ChatGPT Sol 5.6-Extra high under a US government license Extends the number of rows to R_XLEN_T_MAX = 2^52 Edit 2: I just noticed I was the original submitter of this PR from R-forge.
https://github.com/Rdatatable/data.table/blob/master/.github/CONTRIBUTING.md#important-ai-policy-
Proposed patch against commit a3eeb4c
data.table-long-vector-64bit.patch This is made by ChatGPT Sol 5.6-Extra high under a US government license Extends the number of rows to R_XLEN_T_MAX = 2^52 Edit 2: I just noticed I was the original submitter of this PR from R-forge.https://github.com/Rdatatable/data.table/blob/master/.github/CONTRIBUTING.md#important-ai-policy-
https://openai.com/policies/services-agreement/
Especially point 4.1:
4.1. Generally. Customer and Customer’s End Users may provide Input and receive Output. As between Customer and OpenAI, to the extent permitted by applicable law, Customer: (a) retains all ownership rights in Input; and (b) owns all Output. OpenAI hereby assigns to Customer all OpenAI’s right, title, and interest, if any, in and to Output.(Emphasis mine)
Regarding patch.
Did you run our full test suite?
did it succeeded?
is there any runtime overhead that non-long vector are now have to carry as a result of this change?Regarding license and agents.
Could you please double check with your agent that code written was trained only on MPL-2.0 compatible code? ideally let him provide you references.
I asked my agent about that:
https://openai.com/policies/services-agreement/
Especially point 4.1:
4.1. Generally. Customer and Customer’s End Users may provide Input and receive Output. As between Customer and OpenAI, to the extent permitted by applicable law, Customer: (a) retains all ownership rights in Input; and (b) owns all Output. OpenAI hereby assigns to Customer all OpenAI’s right, title, and interest, if any, in and to Output.does it mean if openai write code for me, I can be sure that code produced is compatible license, like GPL or MPL ? answer yes or not.
Received: No.
(Emphasis mine)
Regarding patch. Did you run our full test suite? did it succeeded? is there any runtime overhead that non-long vector are now have to carry as a result of this change?
Regarding license and agents. Could you please double check with your agent that code written was trained only on MPL-2.0 compatible code? ideally let him provide you references.
I asked my agent about that:
https://openai.com/policies/services-agreement/
Especially point 4.1:
4.1. Generally. Customer and Customer’s End Users may provide Input and receive Output. As between Customer and OpenAI, to the extent permitted by applicable law, Customer: (a) retains all ownership rights in Input; and (b) owns all Output. OpenAI hereby assigns to Customer all OpenAI’s right, title, and interest, if any, in and to Output.does it mean if openai write code for me, I can be sure that code produced is compatible license, like GPL or MPL ? answer yes or not.
Received: No.
(Emphasis mine)
I asked ChatGPT Astra High (this time on my own private account):
If a code generated in ChatGPT codex, which is a patch to an MPL licensed open source code, given the services agreement (https://openai.com/policies/services-agreement/ ) and in the US government computer, can it be licensed to one compatible to MPL (to be submitted back to the MPL-licensed project)? This patch was performed as a part of official duty of US government contractors whose contract explicitly mentions that all IPs generated from the contracts belong to the US government. Please answer concisely in 2-3 sentences.Answer: Yes—generally, the patch can be submitted under MPL 2.0 (or the project’s required MPL-compatible contribution terms): OpenAI’s Services Agreement assigns its rights in Codex output to the customer, while MPL permits modifications and requires distributed modifications to covered files to remain under MPL. ([OpenAI]1) Because the contractor agreement expressly assigns work-product IP to the U.S. Government, the Government would ordinarily be the relevant rights holder able to authorize that contribution, subject to the agency’s contract, security, and release policies; Codex output can also carry third-party/open-source licensing obligations that should be checked before submission. ([openai.com]2)
Source: https://chatgpt.com/share/6aa950b5-b75c-83ea-ac09-7dc8134ef77d
About the patch: I ran the full test suite through ChatGPT and it succeeded. I can rerun the suite manually (if you accept the licensing answer; if not, I'm using my own privately and it works)
Overhead: Have not check. Will do. (Subject to an agreeable licensing answer)Dear Roby Joehanes,
Thank you for extending the effort to put a patch together and offer us the code! Currently, the patch has other problems beyond the copyright concerns («Codex output can also carry third-party/open-source licensing obligations that should be checked»), so we cannot accept it. We'll develop a comprehensive solution that will cover all bases that we care about; meanwhile, your fork works well for you. I recommend talking to Håkon Tjeldnes to avoid duplication of effort.
Reacted by Jan Gorecki and AtrebasDear Roby Joehanes,
Thank you for extending the effort to put a patch together and offer us the code! Currently, the patch has other problems beyond the copyright concerns («Codex output can also carry third-party/open-source licensing obligations that should be checked»), so we cannot accept it. We'll develop a comprehensive solution that will cover all bases; meanwhile, your fork works well for you. I recommend talking to Håkon Tjeldnes to avoid duplication of effort.
I did follow up discussion that absolves me from any copyright entanglement: https://chatgpt.com/share/6aa96f6f-00fc-83ea-92e7-42d1fb2b1c96
I see no additional GPL obligation arising from R source code in this patch. The patch appears to be an independently implemented extension of data.table's existing MPL-2.0 code using R's documented/public long-vector and ALTREP APIs, rather than a port of R's GPL implementation.
I would therefore be comfortable describing the output-level provenance as: no identifiable third-party implementation copied from R; R-derived material is limited to public API usage, documented semantics, and generic programming idioms. What cannot truthfully be certified is the maintainer's request that the model was “trained only on MPL-2.0 compatible code”—model training provenance is a separate question and cannot be established from the patch.
Edit: Even Håkon Tjeldnes used Astra xhigh. If you won't accept any AI-made code, then you could not accept his as well.
Just for the record, I am not against using AI.
Asking AI to identify which parts of code need to be updated when I change signature of a function is different than making AI write the change.Our current AI policy could be changed, but that would require data.table team to look into it. Especially Matt Dowle's take on it, as he offered this project to the community.
Allowing AI contribution results in more maintainer time required.
Whoever needs to review or merge PR also needs to understand it, unlike the person who is submitting it, who can produce big and complex patches in minutes.I am a bit with aitap on this issue, even though the new astra model is amazing, and I manually read the tests it makes, I think all of us need a few months to digest what the hell to do with all this magic we can now do, to make some rigorous guidelines etc, the top models can now do more work in 15 minutes than I can do in a month of spare time coding (especially the C code). There is so much potential here for us package developers to improve so much for the whole community, but full human verification is basically impossible.
I guess the way to go would be something similar to the Copilot code review (to get a "verified stamp") combined with some minimum commits / owned repos from the PR author with filters like (not accounts that are < 6 months old etc) to make it more credible.
But for now, we can just work on our forks @robbyjo (+ whoever else want to help us) and when we all feel a bit more comfortable with ai written code we can discuss a possible pull request to the main repo.
Reacted by Michael Chirico and Atrebas
Submitted by: Roby Joehanes; Assigned to: Nobody; R-Forge link
I have a huge table with 4+ billion rows. fread has an integer overflow problem and only read about 73M rows.