Skip to content

More suffixes  #35

Description

@exploy

Hi,

first of all, we really like your parser - it helps us a lot.

Since, we work with quite big sets of data and we found few suffixes which are not included in the config (probably there are more but we just started playing with your parser :) )

If you're interested, we can send more suffixes as we find more :)

Here is a list of suffixes which we added to the config:
'aca', 'acca', 'globally', 'hiring', 'cha', 'cpa', 'csm', 'phr', 'pmp', 'agente'

And here is our test set:
{'full_name': u"Maheen Farooqi, ACA",
'first': u"Maheen",
'last': u"Farooqi"},
{'full_name': u"Howard G. Buffett",
'first': u"Howard",
'last': u"Buffett"},
{'full_name': u"Eunsoo Johanna Jeong",
'first': u"Eunsoo",
'middle': u'Johanna',
'last': u"Jeong"},
{'full_name': u"Ajay Reddy B",
'first': u"Ajay",
'middle': u"Reddy",
'last': u'B'},
{'full_name': u"Kajal Patel Hiring Globally",
'first': u"Kajal",
'last': u'Patel'},
{'full_name': u"Jan Kees de Jager",
'first': u"Jan",
'last': u'de Jager'},
{'full_name': u"Syukri @ Shu Qi Mohd Nor",
'first': u"Syukri",
'last': u'Nor'},
{'full_name': u"Ahmad Jayadi, CHA",
'first': u"Ahmad",
'last': u'Jayadi'},
{'full_name': u"Will De Groot",
'first': u"Will",
'last': u'De Groot'},
{'full_name': u"Sander van 't Noordende",
'first': u"Sander",
'last': u'Noordende'},
{'full_name': u"Dr. Mazen Elrouqi",
'first': u"Mazen",
'last': u'Elrouqi'},
{'full_name': u"Reem Al-Bahrani, CPA",
'first': u"Reem",
'last': u'Al-Bahrani'},
{'full_name': u"Cherice J. (Johnson) Williams",
'first': u"Cherice",
'last': u'Williams'},
{'full_name': u"Eddie Chang Seng Dee Adecco",
'first': u"Eddie",
'last': u'Adecco'},
{'full_name': u"Hisham Ibrahim, CPA, CFA",
'first': u"Hisham",
'last': u'Ibrahim'},
{'full_name': u"Lamha Yahya Al Mawali",
'first': u"Lamha",
'last': u'Mawali'},
{'full_name': u"Lisa Schmidt, MBA",
'first': u"Lisa",
'last': u'Schmidt'},
{'full_name': u"Sejal Chaturvedi, CSM",
'first': u"Sejal",
'last': u'Chaturvedi'},
{'full_name': u"Arivendhi Pargunan, MBA, PMP, CMQ/OE",
'first': u"Arivendhi",
'last': u'Pargunan'},
{'full_name': u"Maura Cellario-Pingor, PHR",
'first': u"Maura",
'last': u'Cellario-Pingor'},
{'full_name': u"Gary Strawsburg, PMP, CSCP, CLSSBB",
'first': u"Gary",
'last': u'Strawsburg'},
{'full_name': u"Nicolo' Pantaleo Agente",
'first': u"Nicolo'",
'last': u'Pantaleo'}

Regards,
Jarek

Activity

  1. derek73 commented on Aug 4, 2015

    @derek73
    Owner

    Hi Jarek,

    Thanks for the compliments and the additional surnames.

    To add new surnames to the project's config, we just have to make sure that they are not likely to also be last names. Any word that is in the suffix set will never be considered a last name by the parser.

    It looks like these could potentially be last names:

    (Sometimes it seems like any combination of letters that includes vowels is a potentially name in some language.) Most of those last names look pretty darn rare. Though, maybe they are equally as rare as suffixes? It's hard to tell what would be correct more often.

    For "Kajal Patel Hiring Globally", it seems like "Hiring Globally" is not really a part of their name, right? It's a company name? We probably don't want to get in the business of adding company names to the project's suffixes config or we'd have a very large list. That's probably better suited for some kind of machine learning version of this library.

    Also, I'm curious, what is your use case for this library? What can you tell me about your dataset? I get very little feedback on how people are actually using it, so it's hard to know if it's structured in a helpful way or if there are things that would make it easier to use in your own code. Any feedback you have would be appreciated. Also, depending on where your dataset is coming from, you may have a better handle on what's common than I do.

  2. added this to the v0.3.5 milestone on Aug 4, 2015
  3. self-assigned this
    on Aug 4, 2015
  4. exploy commented on Aug 4, 2015

    @exploy
    Author

    Hi Derek,

    many thanks for an extensive answer. It forced me to look, on my case, from a different angle.

    What do you think about splitting those full phrases (which I sent last time) by first comma and analyze only first part? It works for majority of the cases for us.

    Our case is to extract first, middle and last name from the full name of a person on LinkedIn public profile - here is an example profile https://www.linkedin.com/in/cindylaughlin.

    Let us know if we can help you more with development.

    Regards,
    Jarek

  5. derek73 commented on Aug 4, 2015

    @derek73
    Owner

    I'm not sure I understand your question about splitting the full phrase by commas. The parser supports 3 different comma formats for names, noted in the read me.

    If your dataset does not contain the "lastname, firstname" format, you could probably do like you suggest in pre-processing. Just split your string on commas and run the first part, without commas, through the parser, then append the other comma pieces that you got in preprocessing to hn_instance.titles_list (or wherever you want them) and everything should work the same.

    I added a few of those safer titles in v0.3.5, just uploaded to pypi.

  6. exploy commented on Aug 4, 2015

    @exploy
    Author

    That's exactly what I meant and I implemented exactly the suggestion you made :)

    Thanks for help and all the best,
    Jarek

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Projects

No projects

    Milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions