Repository navigation
datatable.prettyprint.char could use ellipsis (…) instead of three dots #7715
Description
Activity
related to #5425
Not sure using ellipsis is the best option here, since locales without UTF-8 support will then display strange characters.
Also AFAIK ellipsis is a double-width character in certain Asian fonts and this might break alignment of columns?
those are valid points. what locales, specifically? can you please provide a code example?
Things that should not work are ASCII or latin1.
iconv("\u2026", "UTF-8", "ASCII") # [1] NA iconv("\u2026", "UTF-8", "latin1") # [1] NA
I guess we could check with
l10n_info()$'UTF-8'and fall back to three dots.Reacted by Toby Dylan Hocking and Tomasz JanczykAlso ccing @aitap here as he knows more about encodings
Hello @ben-schwen @tdhock , thank you for your remarks.
Not sure using ellipsis is the best option here, since locales without UTF-8 support will then display strange characters.
Also AFAIK ellipsis is a double-width character in certain Asian fonts and this might break alignment of columns?
I updated the related PR so it falls back to three dots when locale is not UTF-8, seems to be working with locales that don't support it, e.g:
Sys.setlocale("LC_CTYPE", "Chinese") > l10n_info()$`UTF-8` [1] FALSEoptions(datatable.prettyprint.char=2, datatable.print.class=FALSE) data.table::data.table(x="foo") x 1: fo...Do you have any suggestions how the added
ifclause should be covered in unit tests? Currently test2253*is run only for UTF-8 locale. Should any tests be added?
Thanks!Reacted by Toby Dylan Hocking and aitapI was surprised to learn (by
greppinglibiconvsource code for0x2026) that there are many single-byte encodings in use that include…(e.g.0x85in at least some ANSI encodings on Windows). The test should probably be foridentical(enc2native("\u2026"), "\u2026"), for which we have theutf8_checkhelper. I'm afraid the tests for abbreviated output will have to become conditional on that as well.Reacted by Toby Dylan Hockingusing #7788 I see up to six dots in a row. three dots from list abbreviation, then another three from prettyprint
> options(width = 17, datatable.prettyprint.char = NULL);data.table(x = "0123456789", L = list(1:25)) x <char> 1: 0123456789 L <list> 1: 1,2,3,4,5,6... > options(width = 18, datatable.prettyprint.char = NULL);data.table(x = "0123456789", L = list(1:25)) x <char> 1: 0123456789 L <list> 1: 1,2,3,4,5,6,... > options(width = 19, datatable.prettyprint.char = NULL);data.table(x = "0123456789", L = list(1:25)) x <char> 1: 0123456789 L <list> 1: 1,2,3,4,5,6,.... > options(width = 20, datatable.prettyprint.char = NULL);data.table(x = "0123456789", L = list(1:25)) x <char> 1: 0123456789 L <list> 1: 1,2,3,4,5,6,..... > options(width = 21, datatable.prettyprint.char = NULL);data.table(x = "0123456789", L = list(1:25)) x <char> 1: 0123456789 L <list> 1: 1,2,3,4,5,6,...... > options(width = 22, datatable.prettyprint.char = NULL);data.table(x = "0123456789", L = list(1:25)) x <char> 1: 0123456789 L <list> 1: 1,2,3,4,5,6,...[... > options(width = 23, datatable.prettyprint.char = NULL);data.table(x = "0123456789", L = list(1:25)) x <char> 1: 0123456789 L <list> 1: 1,2,3,4,5,6,...[2... > options(width = 24, datatable.prettyprint.char = NULL);data.table(x = "0123456789", L = list(1:25)) x <char> 1: 0123456789 L <list> 1: 1,2,3,4,5,6,...[25... > options(width = 25, datatable.prettyprint.char = NULL);data.table(x = "0123456789", L = list(1:25)) x <char> 1: 0123456789 L <list> 1: 1,2,3,4,5,6,...[25]
this is a silly trivial example, but
The first output above is actually wider than the second. (5 characters instead of 3, surprising, because it is supposed to abbreviate)
could we change to below output? (still three characters)