Repository navigation
Multiple small issues with messages in the C code #6503
Description
Activity
- addedtranslationissues/PRs related to message translation projectsissues/PRs related to message translation projects
on Sep 17, 2024 Should this mention the issue tracker instead?
yes
Reacted by aitapThis one is actually possible for me to translate in parts (so we can keep it as is if needed)
I guess you're referring to
few/many, not thejump>0complete sentence added at the end. If so: yes, I think we should branch it.I think we shouldn't be calling it ASCII if its / is not immediately before 0. How about "The character '/' is not just before '0' in the source character set
Hmm, good point, though I worry "source character set" is a highly technical term --> relatively tough to translate. WDYT about "Unlike the very common case, e.g. ASCII, the character '/' is not just before '0'".
Would it be more helpful to mention character strings in the error message too?
Ping @jangorecki who has the best context here
Is LENGTH(.) < 0 something that's possible in R ≥ 3.3.0?
Good spot. Has it ever been possible? My best guess was this was a somewhat half-baked fix here:
Author fixed the
length=0case to work and changed the message to adapt without stopping to think "length<0is not possible" instead of "how do I adaptx<=0 | x>0to includex=0valid?x<0 | x>=0".So I think we can just drop that condition and the corresponding part of the message.
I guess you're referring to
few/many, not thejump>0complete sentence added at the end. If so: yes, I think we should branch it.Yes, I mean the part about
few/many. Could you please clarify what you meant by "branch it"?I worry "source character set" is a highly technical term --> relatively tough to translate. WDYT about "Unlike the very common case, e.g. ASCII, the character '/' is not just before '0'".
Agreed, this phrasing sounds fine.
My best guess was this was a somewhat half-baked fix here:
...which started out as a check for "opposite of
length() > 0". Will remove the condition and adjust the message.Could you please clarify what you meant by "branch it"?
thisNcol < ncol ? _("A line with too few fields...") : _("A line with too many fields...)"Reacted by aitapNo idea really, but I can imagine some hacks in setting up length to negative value, that could have been in play.
some contribs
- should not be translatated:
Line 1343 in 3b376da
DTPRINT(_(" NAstrings = [")); - messages like this are often hard to translate
Lines 224 to 225 in 3b376da
if (verbose) Rprintf(_("gather took ... ")); switch (TYPEOF(x)) { - Better to write
STOP(_("Internal error in %s: %s. Please report to the data.table issues tracker", __func__, internalErr);
Line 2627 in 3b376da
STOP("%s %s: %s. %s", _("Internal error in"), __func__, internalErr, _("Please report to the data.table issues tracker")); // # nocov - part of above message, perhaps difficult to translate out of context.
Line 1711 in 3b376da
DTPRINT(_(" with %d fields using quote rule %d\n"), topNumFields, quoteRule);
Reacted by aitap- should not be translatated:
[aside/FYI @rikivillalba if your permalink includes column numbers (
C10,C23in your 2nd bullet before my edit), it won't render inline in the issue. I don't know why GitHub sometimes includes the column numbers]Reacted by Ricardo Villalba@jangorecki I'm not seeing calls to
SETLENGTHinfmelt.c, but I do know thatdata.tableuses negativeTRUELENGTHfor its internal purposes. Hmm.@rikivillalba Thank you for reminding me! There's quite a lot of debugging printout that consists of argument/variable names that (arguably) should not be translated:
-
Line 440 in d7e95d1
if (verbose) Rprintf(_("RHS_list_of_columns == %s\n"), RHS_list_of_columns ? "true" : "false"); -
Line 69 in d7e95d1
if (iN && it!=xt) error(_("typeof x.%s (%s) != typeof i.%s (%s)"), CHAR(STRING_ELT(getAttrib(xdt,R_NamesSymbol),xcols[col]-1)), type2char(xt), CHAR(STRING_ELT(getAttrib(idt,R_NamesSymbol),icols[col]-1)), type2char(it)); -
Line 109 in d7e95d1
if (isNull(jiscols) && (length(bynames)!=length(groups) || length(bynames)!=length(grpcols))) error(_("!length(bynames)[%d]==length(groups)[%d]==length(grpcols)[%d]"),length(bynames),length(groups),length(grpcols)); -
Line 138 in d7e95d1
if (length(names) != length(SDall)) error(_("length(names)!=length(SD)")); -
Line 154 in d7e95d1
if (length(xknames) != length(xSD)) error(_("length(xknames)!=length(xSD)")); -
Lines 162 to 163 in d7e95d1
if (length(iSD)!=length(jiscols)) error(_("length(iSD)[%d] != length(jiscols)[%d]"),length(iSD),length(jiscols)); if (length(xSD)!=length(xjiscols)) error(_("length(xSD)[%d] != length(xjiscols)[%d]"),length(xSD),length(xjiscols)); -
Line 402 in d7e95d1
if (verbose) Rprintf(_("dogroups: growing from %d to %d rows\n"), length(VECTOR_ELT(ans,0)), estn); -
Line 784 in d7e95d1
Rprintf(_("nradix=%d\n"), nradix); - One more fragmented sentence, mixing an untranslatable variable name (
jump0size==0) in the middle:
Lines 1891 to 1896 in d7e95d1
DTPRINT(_(" Number of sampling jump points = %d because "), nJumps); if (nrowLimit<INT64_MAX) DTPRINT(_("nrow limit (%"PRIu64") supplied\n"), (uint64_t)nrowLimit); else if (jump0size==0) DTPRINT(_("jump0size==0\n")); else DTPRINT(_("(%"PRIu64" bytes from row 1 to eof) / (2 * %"PRIu64" jump0size) == %"PRIu64"\n"), (uint64_t)sz, (uint64_t)jump0size, (uint64_t)(sz/(2*jump0size))); } -
Line 2261 in d7e95d1
if (verbose) DTPRINT(_(" jumps=[%d..%d), chunk_size=%"PRIu64", total_size=%"PRIu64"\n"), -
Line 158 in d7e95d1
if (INTEGER(nThreadArg)[0]<1) error(_("nThread(%d)<1"), INTEGER(nThreadArg)[0]); -
Line 122 in d7e95d1
if (verbose) Rprintf(_("nth=%d, nBatch=%d\n"),nth,nBatch); -
Line 178 in d7e95d1
if (verbose) Rprintf(_("maxBit=%d; MSBNbits=%d; shift=%d; MSBsize=%zu\n"), maxBit, MSBNbits, shift, MSBsize); -
Lines 638 to 639 in d7e95d1
DTPRINT(_("\nargs.doRowNames=%d args.rowNames=%p args.rowNameFun=%d doQuote=%d args.nrow=%"PRId64" args.ncol=%d eolLen=%d\n"), args.doRowNames, args.rowNames, args.rowNameFun, doQuote, args.nrow, args.ncol, eolLen); -
Line 63 in d7e95d1
if (LENGTH(f) != ngrp) error(_("length(f)=%d != length(l)=%d"), LENGTH(f), ngrp); - This one uses untranslated default values in
mygetenv, which may be not that important:
Lines 89 to 97 in d7e95d1
Rprintf(_(" omp_get_num_procs() %d\n"), omp_get_num_procs()); Rprintf(_(" R_DATATABLE_NUM_PROCS_PERCENT %s\n"), mygetenv("R_DATATABLE_NUM_PROCS_PERCENT", "unset (default 50)")); Rprintf(_(" R_DATATABLE_NUM_THREADS %s\n"), mygetenv("R_DATATABLE_NUM_THREADS", "unset")); Rprintf(_(" R_DATATABLE_THROTTLE %s\n"), mygetenv("R_DATATABLE_THROTTLE", "unset (default 1024)")); Rprintf(_(" omp_get_thread_limit() %d\n"), omp_get_thread_limit()); Rprintf(_(" omp_get_max_threads() %d\n"), omp_get_max_threads()); Rprintf(_(" OMP_THREAD_LIMIT %s\n"), mygetenv("OMP_THREAD_LIMIT", "unset")); // CRAN sets to 2 Rprintf(_(" OMP_NUM_THREADS %s\n"), mygetenv("OMP_NUM_THREADS", "unset")); Rprintf(_(" RestoreAfterFork %s\n"), RestoreAfterFork ? "true" : "false"); -
Line 39 in d7e95d1
if (length(order) != nrow) error(_("nrow(x)[%d]!=length(order)[%d]"),nrow,length(order));
Which of these should have their
_removed? Which should be translated anyway?-
for those, I think we need a system for marking "notranslate" at the same time, so that CI/release doesn't keep finding these strings that are in translateable calls, just not useful to translate.
In potools, that's
// # notranslate.Let's open another separate issue for those, it's a bit different from the topic at hand.
Reacted by aitap- This one indicates a problem with
xgettext: the string eventually given togettext()will not be translated becausexgettexthad split it into parts, including"%":
Lines 944 to 945 in d7e95d1
? _("%"FMT" (type '%s') at RHS position %d "TO" when assigning to type '%s' (target vector)") \ : _("%"FMT" (type '%s') at RHS position %d "TO" when assigning to type '%s' (column %d named '%s')"), \
Edit: hmm, I don't see "at RHS position" at all in the official
*.potfile. Nevertheless, I don't thinkxgettexthandles compile-time string concatenation in general.Reacted by Michael Chirico- This one indicates a problem with
Re: the previous comment, any suggested fix? The first things that come to mind seem messy. Maybe we should functionalize the macro...
The core of the problem is that even if we do follow the recommendation to format the number into a temporary buffer and then use
%sfor all number formats, we still have a combinatorial explosion of 7 (TO) ⋅ 2 (target vector/column %d named '%s') strings thatxgettextwants spelled out in the source code.Can we cheat a little and split the clauses into two sentences? Then we'll have
_("%s (type '%s') at RHS position %d taken as TRUE") _("%s (type '%s') at RHS position %d taken as 0") _("%s (type '%s') at RHS position %d either truncated (precision lost) or taken as 0") _("%s (type '%s') at RHS position %d out-of-range (NA)") _("%s (type '%s') at RHS position %d out-of-range (NA) or truncated (precision lost)") _("%s (type '%s') at RHS position %d had either imaginary part discarded or real part truncated (precision lost)") _("%s (type '%s') at RHS position %d had imaginary part discarded")and
_("Problem when assigning to type '%s' (target vector):") _("Problem when assigning to type '%s' (column %d named '%s'):")...which is more manageable. The strings could be rephrased further, making them truly separate sentences (
The problem happened when assigning to type '%s' (target vector).).Would it be more helpful to mention character strings in the error message too?
cc @jangorecki. I'm not sure what change you have in mind exactly, but I think the call-out to consult
?substitute2is the most helpful part. I worry adding to the message will make it over-complicated.part of above message, perhaps difficult to translate out of context.
Good spot, there's a similar fragment below. It's a bit hard to pull apart exactly how we'd make this friendlier. I think the best bet is to make each fragment a standalone sentence. I'll want to read more carefully how exactly those show up in the verbose output. Filing as a separate issue so it's not lost, there's already a lot going on here.
This is a continuation of #6482. I think I won't find any more issues in the C code messages. Some of these are questions, a few indicate real translation hurdles.
data.table/src/fread.c
Lines 1933 to 1934 in 9443409
Translation in parts could be done reliably with
pgettext, which seems to be currently unavailable even to C code in R's built-in copy ofgettexton Windows.data.table/src/fsort.c
Lines 255 to 259 in 9443409
_("no")not to mean something else in a different part of the code:data.table/src/fread.c
Line 2038 in 9443409
Unfortunately, expanding this sentence will duplicate the full line of the code.
The rest are questions unrelated to translations:
/is not immediately before0. How about "The character '/' is not just before '0' in the source character set"?data.table/src/init.c
Line 225 in 9443409
data.table/src/init.c
Line 227 in 9443409
.Last.updated:data.table/src/init.c
Line 367 in 9443409
data.table:::list2langconverts character strings to symbols. Would it be more helpful to mention character strings in the error message too?data.table/src/programming.c
Line 16 in 9443409
data.table/src/fmelt.c
Line 805 in 9443409
LENGTH(.) < 0something that's possible in R ≥ 3.3.0?data.table/src/fmelt.c
Line 66 in 9443409
(Checkboxes indicate either no action needed or a corresponding fix being suggested in #6504)