Repository navigation
Can't handle Japanese names #34
Description
Activity
Yes, the name parser is really only relevant for common names in latin-based languages.
It does handle unicode though. The error you posted is telling you that those characters cannot be encoded using ascii, which is correct. Likely you need are using python 2.x and need to add this to the first line of your file so it doesn't assume you are using ascii:
# -*- coding: utf-8 -*-I'm curious, what would you like the parser to do with Chinese names? Are there parts of a Chinese name that would be helpful for a parser to separate?
I just wanted to point out that "last name" and "first name" are rather
Euro-centric ideas — "family name" or "surname" and "given name" or
"personal name" might have been better choices. But probably too late now
to change your API.The example I gave was Japanese, not Chinese. Wikipedia has good articles
on Chinese and Japanese names, which might interest you. Chinese names are
a nightmare to extract from text, but are relatively easy to handle once
extracted because they have a relatively small number of common family
names (almost all are a single character (汉字 or 漢字)) and most given names
are two characters. Japanese, on the other hand, have a large number of
family names; and both family and given names vary in length. Both Chinese
and Japanese do not put a space between family and personal names, so
separating the two is a non-trivial exercise (much more difficult with
Japanese). They also have complicated "titles" that come after the name
(さん、様、殿、氏、先生、etc., etc.). I really wouldn't expect you to get it right, but
it's unfortunate that you seem to have chosen an API the makes some
unnecessary presumptions if someone in the future decides to tackle them.
[I don't know enough about Korean, Thai, Indonesian, Hindi, Tamil, etc.
names to say anything about them; but some of these are written with
Latin-1 and don't fit with your design either.]As to the unicode problem -- I'm not sure how things are set up on my Mac;
but between Emacs, Mac, and Python 2.7.6 anything is possible and a bit of
playing with encode("utf-8") and decode("utf-8") together with # coding:
utf-8 didn't fix it.Cheers,
- peter
.
On 27 July 2015 at 22:26, derek73 [email protected] wrote:
Yes, the name parser is really only relevant for common names in
latin-based languages.It does handle unicode though. The error you posted is telling you that
those characters cannot be decoded using ascii, which is correct. Likely
you need are using python 2.x and need to add this to the first line of
your file so it doesn't assume you are using ascii:-- coding: utf-8 --
I'm curious, what would you like the parser to do with Chinese names? Are
there parts of a Chinese name that would be helpful for a parser to
separate?—
Reply to this email directly or view it on GitHub
#34 (comment)
.- peter
re: unicode, you could take a look at the project's tests. They include unicode in the file and do not throw unicode errors. Might help you discover where the problem is.
This parser's logic starts by splitting a string on spaces, which sounds like would be completely useless for Chinese and Japanese names. So, it seems like there's nothing terribly helpful that we could easily do to this parser to make it more useful for those languages. Going to close out this ticket. Feel free to reopen if you have other suggestions.
- added a commit that references this issue
on Aug 4, 2015 I looked into this a bit more and noticed 2 things. You don't have the little unicode marker in your example string.. Perhaps you need:
name = nameparser.HumanName(u"鈴木太郎")or to work in both python 2.x & 3.x:
from __future__ import unicode_literals name = nameparser.HumanName("鈴木太郎")I also did some more poking around and noticed that the string methods that return bytes were not being encoded correctly in python 2.x. It might have caused the unicode problem you spotted if especially if you were playing around in the terminal. If your terminal is not using utf8 encoding you might need to do something like:
HumanName(name, encoding=sys.stdout.encoding)fixes are in v0.3.5 now
Also: the concept of "last name" and "first name" isn't valid for Chinese, Japanese, Korean (CJK) names.