Repository navigation
[Request] fread to support thousand separator #1636
Description
Activity
- changed the title
[-]Can `fread` support thousand separator?[/-][+][Request] `fread` to support thousand separator[/+]on Jul 4, 2017 NB: This is likely to result in big performance improvements because of the string cache issue we've encountered a few other places -- if the thousand separator is recognized, those columns will be read as numeric, not character.
Reacted by Egor Kotov, Nitish Jha, Floris Padt, Mukul, Michael Chirico and Michael Mayer- addedtop requestOne of our most-requested issuesOne of our most-requested issues
on Apr 14, 2024 Hi @MichaelChirico ,
I have reviewed the issue and I believe I can help resolve it. Before proceeding, I wanted to check with you if it would be okay for me to work on this issue.Here is my proposed approach to implement thousand separator feature in fread:
To implement the thousands separator support in fread(), we first modify the R interface by adding a thousands argument, defaulting to an empty string ("") for clarity. This argument is validated to ensure it is a single non-numeric character, distinct from both the decimal separator (dec) and field separator (sep). Once validated, it is passed to C through .Call(), allowing seamless integration into the parsing logic.
In C, we extend the FreadGlobalArgs struct by adding a thousands_sep field to store the separator globally, ensuring consistent access throughout the parsing process. To handle separator stripping efficiently within multi-threaded execution, we introduce a thread-local buffer (thousands_stripping_buffer) inside FreadThreadLocal. This design maintains thread safety and prevents race conditions, as each thread works with its own independent buffer.
To optimize memory usage and avoid frequent allocations, each thread initializes a 1024-byte buffer when it starts. This size is chosen as a reasonable default to accommodate most numeric fields without excessive memory overhead. The buffer expands dynamically when needed, ensuring efficiency without excessive reallocation. Additionally, memory is properly freed during thread cleanup to prevent leaks, maintaining stability in long-running processes.
For parsing, we first detect and validate whether a numeric field contains a valid thousands separator. The separator must be correctly placed—such as "1,000" being valid but "1,,000" or ",1000" being invalid. If valid, we strip the separators only when necessary, using the pre-allocated thread-local buffer instead of allocating new memory for each field. This reduces unnecessary processing and improves performance.
To ensure robustness, improperly formatted numbers (e.g., "1,,000", ",1000") are either converted to character columns or assigned NA, depending on overall column consistency. I'll also take care to correctly handle different combinations of thousands and dec separators (e.g., thousands="," with dec="."), preventing misinterpretation of numeric values. This approach balances correctness and help in improving performance.
Please let me know if you have any feedback or suggestions. I look forward to your response.
Best regards,
MukulIn C, we extend the FreadGlobalArgs struct by adding a thousands_sep field to store the separator globally, ensuring consistent access throughout the parsing process. To handle separator stripping efficiently within multi-threaded execution, we introduce a thread-local buffer (thousands_stripping_buffer) inside FreadThreadLocal. This design maintains thread safety and prevents race conditions, as each thread works with its own independent buffer.
This part doesn't sound right to me. See #4482.
More generally I'm not sure the right interface here.
- AIUI "thousands" is a locale-specific thing. Chinese and Japanese might use 4-digit (myriad) separators, i.e. 1e6 written as 100,0000, 1e10 written as 100,0000,0000, etc. Do we want to restrict ourselves to only supporting the 3-digit case? https://en.wikipedia.org/wiki/Decimal_separator#Digit_grouping
- Do we want to offer automatic detection of the separator? How? It seems likely to cause major conflicts with the very common case of
sep=",". It may indeed be a breaking change -- what will happen in files likea,b,c\n123,456,789\nwhere there's potential ambiguity on whether certain sub-regions of the file are 3- vs. more-digit numbers? Per the above Wikipedia article, there's also a more general concern ofsep=conflicting with the proposedthousands=. In principle a file could be well-formed like so:a,b,c\n"123,456","45,678","56,789"\n, how will our automaticsep=detection interact with that? - How will this interact with
sep2? when will sep2 in fread be implemented? #1162
improperly formatted numbers... are either converted to character columns or assigned NA
We should behave consistently to the prior behavior -- how does
fread()treat improperly formatted numbers in columns it assumes are numeric today? IINM it will actually drop the assumption the field is numeric and re-read the column as character. This is to ensure no information is lost byfread().Overall, this feature is likely to be quite messy to implement as it has a lot of interactions with existing features. I might recommend working on #1162 instead, which should be less fraught with these concerns.
Reacted by Mukul
I wonder if
freadcould gain a argument to recognize thousand separator like in pandas