Skip to content

Can't handle Japanese names #34

Description

@pludemann

name = nameparser.HumanName("鈴木太郎")

name
Traceback (most recent call last):
File "", line 1, in
UnicodeEncodeError: 'ascii' codec can't encode characters in position 36-39: ordinal not in range(128)

Also: the concept of "last name" and "first name" isn't valid for Chinese, Japanese, Korean (CJK) names.

Activity

  1. derek73 commented on Jul 28, 2015

    @derek73
    Owner

    Yes, the name parser is really only relevant for common names in latin-based languages.

    It does handle unicode though. The error you posted is telling you that those characters cannot be encoded using ascii, which is correct. Likely you need are using python 2.x and need to add this to the first line of your file so it doesn't assume you are using ascii:

    # -*- coding: utf-8 -*-
    

    I'm curious, what would you like the parser to do with Chinese names? Are there parts of a Chinese name that would be helpful for a parser to separate?

  2. pludemann commented on Jul 28, 2015

    @pludemann
    Author

    I just wanted to point out that "last name" and "first name" are rather
    Euro-centric ideas — "family name" or "surname" and "given name" or
    "personal name" might have been better choices. But probably too late now
    to change your API.

    The example I gave was Japanese, not Chinese. Wikipedia has good articles
    on Chinese and Japanese names, which might interest you. Chinese names are
    a nightmare to extract from text, but are relatively easy to handle once
    extracted because they have a relatively small number of common family
    names (almost all are a single character (汉字 or 漢字)) and most given names
    are two characters. Japanese, on the other hand, have a large number of
    family names; and both family and given names vary in length. Both Chinese
    and Japanese do not put a space between family and personal names, so
    separating the two is a non-trivial exercise (much more difficult with
    Japanese). They also have complicated "titles" that come after the name
    (さん、様、殿、氏、先生、etc., etc.). I really wouldn't expect you to get it right, but
    it's unfortunate that you seem to have chosen an API the makes some
    unnecessary presumptions if someone in the future decides to tackle them.
    [I don't know enough about Korean, Thai, Indonesian, Hindi, Tamil, etc.
    names to say anything about them; but some of these are written with
    Latin-1 and don't fit with your design either.]

    As to the unicode problem -- I'm not sure how things are set up on my Mac;
    but between Emacs, Mac, and Python 2.7.6 anything is possible and a bit of
    playing with encode("utf-8") and decode("utf-8") together with # coding:
    utf-8 didn't fix it.

    Cheers,

    • peter
      .

    On 27 July 2015 at 22:26, derek73 [email protected] wrote:

    Yes, the name parser is really only relevant for common names in
    latin-based languages.

    It does handle unicode though. The error you posted is telling you that
    those characters cannot be decoded using ascii, which is correct. Likely
    you need are using python 2.x and need to add this to the first line of
    your file so it doesn't assume you are using ascii:

    -- coding: utf-8 --

    I'm curious, what would you like the parser to do with Chinese names? Are
    there parts of a Chinese name that would be helpful for a parser to
    separate?

    —
    Reply to this email directly or view it on GitHub
    #34 (comment)
    .

  3. derek73 commented on Aug 4, 2015

    @derek73
    Owner

    re: unicode, you could take a look at the project's tests. They include unicode in the file and do not throw unicode errors. Might help you discover where the problem is.

    This parser's logic starts by splitting a string on spaces, which sounds like would be completely useless for Chinese and Japanese names. So, it seems like there's nothing terribly helpful that we could easily do to this parser to make it more useful for those languages. Going to close out this ticket. Feel free to reopen if you have other suggestions.

  4. derek73 commented on Aug 4, 2015

    @derek73
    Owner

    I looked into this a bit more and noticed 2 things. You don't have the little unicode marker in your example string.. Perhaps you need:

    name = nameparser.HumanName(u"鈴木太郎")
    

    or to work in both python 2.x & 3.x:

    from __future__ import unicode_literals
    name = nameparser.HumanName("鈴木太郎")
    

    I also did some more poking around and noticed that the string methods that return bytes were not being encoded correctly in python 2.x. It might have caused the unicode problem you spotted if especially if you were playing around in the terminal. If your terminal is not using utf8 encoding you might need to do something like:

    HumanName(name, encoding=sys.stdout.encoding)

  5. reopened this on Aug 4, 2015
  6. added this to the v0.3.5 milestone on Aug 4, 2015
  7. self-assigned this
    on Aug 4, 2015
  8. derek73 commented on Aug 4, 2015

    @derek73
    Owner

    fixes are in v0.3.5 now

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Projects

No projects

    Milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions