Skip to content

[Request] fread to support thousand separator #1636

Description

@ywhcuhk

I wonder if fread could gain a argument to recognize thousand separator like in pandas

Activity

  1. changed the title [-]Can `fread` support thousand separator?[/-] [+][Request] `fread` to support thousand separator[/+] on Jul 4, 2017
  2. MichaelChirico commented on Apr 12, 2024

    @MichaelChirico
    Member

    NB: This is likely to result in big performance improvements because of the string cache issue we've encountered a few other places -- if the thousand separator is recognized, those columns will be read as numeric, not character.

  3. Mukulyadav2004 commented on Apr 1, 2025

    @Mukulyadav2004
    Contributor

    Hi @MichaelChirico ,
    I have reviewed the issue and I believe I can help resolve it. Before proceeding, I wanted to check with you if it would be okay for me to work on this issue.

    Here is my proposed approach to implement thousand separator feature in fread:

    To implement the thousands separator support in fread(), we first modify the R interface by adding a thousands argument, defaulting to an empty string ("") for clarity. This argument is validated to ensure it is a single non-numeric character, distinct from both the decimal separator (dec) and field separator (sep). Once validated, it is passed to C through .Call(), allowing seamless integration into the parsing logic.

    In C, we extend the FreadGlobalArgs struct by adding a thousands_sep field to store the separator globally, ensuring consistent access throughout the parsing process. To handle separator stripping efficiently within multi-threaded execution, we introduce a thread-local buffer (thousands_stripping_buffer) inside FreadThreadLocal. This design maintains thread safety and prevents race conditions, as each thread works with its own independent buffer.

    To optimize memory usage and avoid frequent allocations, each thread initializes a 1024-byte buffer when it starts. This size is chosen as a reasonable default to accommodate most numeric fields without excessive memory overhead. The buffer expands dynamically when needed, ensuring efficiency without excessive reallocation. Additionally, memory is properly freed during thread cleanup to prevent leaks, maintaining stability in long-running processes.

    For parsing, we first detect and validate whether a numeric field contains a valid thousands separator. The separator must be correctly placed—such as "1,000" being valid but "1,,000" or ",1000" being invalid. If valid, we strip the separators only when necessary, using the pre-allocated thread-local buffer instead of allocating new memory for each field. This reduces unnecessary processing and improves performance.

    To ensure robustness, improperly formatted numbers (e.g., "1,,000", ",1000") are either converted to character columns or assigned NA, depending on overall column consistency. I'll also take care to correctly handle different combinations of thousands and dec separators (e.g., thousands="," with dec="."), preventing misinterpretation of numeric values. This approach balances correctness and help in improving performance.

    Please let me know if you have any feedback or suggestions. I look forward to your response.

    Best regards,
    Mukul

  4. MichaelChirico commented on Apr 1, 2025

    @MichaelChirico
    Member

    In C, we extend the FreadGlobalArgs struct by adding a thousands_sep field to store the separator globally, ensuring consistent access throughout the parsing process. To handle separator stripping efficiently within multi-threaded execution, we introduce a thread-local buffer (thousands_stripping_buffer) inside FreadThreadLocal. This design maintains thread safety and prevents race conditions, as each thread works with its own independent buffer.

    This part doesn't sound right to me. See #4482.

    More generally I'm not sure the right interface here.

    • AIUI "thousands" is a locale-specific thing. Chinese and Japanese might use 4-digit (myriad) separators, i.e. 1e6 written as 100,0000, 1e10 written as 100,0000,0000, etc. Do we want to restrict ourselves to only supporting the 3-digit case? https://en.wikipedia.org/wiki/Decimal_separator#Digit_grouping
    • Do we want to offer automatic detection of the separator? How? It seems likely to cause major conflicts with the very common case of sep=",". It may indeed be a breaking change -- what will happen in files like a,b,c\n123,456,789\n where there's potential ambiguity on whether certain sub-regions of the file are 3- vs. more-digit numbers? Per the above Wikipedia article, there's also a more general concern of sep= conflicting with the proposed thousands=. In principle a file could be well-formed like so: a,b,c\n"123,456","45,678","56,789"\n, how will our automatic sep= detection interact with that?
    • How will this interact with sep2? when will sep2 in fread be implemented? #1162

    improperly formatted numbers... are either converted to character columns or assigned NA

    We should behave consistently to the prior behavior -- how does fread() treat improperly formatted numbers in columns it assumes are numeric today? IINM it will actually drop the assumption the field is numeric and re-read the column as character. This is to ensure no information is lost by fread().

    Overall, this feature is likely to be quite messy to implement as it has a lot of interactions with existing features. I might recommend working on #1162 instead, which should be less fraught with these concerns.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions