Supporting Employer Name Normalization at both Entity and Cluster Level

In the recruitment domain, the employer name normalization task, which links employer names in job postings or resumes to entities in an employer knowledge base (KB), is important to many business applications. In previous work, we proposed the CompanyDepot system, which used machine learning techniques to address the problem. After applying it to several applications at CareerBuilder, we faced several new challenges: 1) how to avoid duplicate normalization results when the KB is noisy and contains many duplicate entities; 2) how to address the vocabulary gap between query names and entity names in the KB; and 3) how to use the context available in jobs and resumes to improve normalization quality. To address these challenges, in this paper we extend the previous CompanyDepot system to normalize employer names not only at entity level, but also at cluster level by mapping a query to a cluster in the KB that best matches the query. We also propose a new metric based on success rate and diversity reduction ratio for evaluating the cluster-level normalization. Moreover, we perform query expansion based on five data sources to address the vocabulary gap challenge and leverage the url context for the employer names in many jobs and resumes to improve normalization quality. We show that the proposed CompanyDepot-V2 system outperforms the previous CompanyDepot system and several other baseline systems over multiple real-world datasets. We also demonstrate the large improvement on normalization quality from entity-level to cluster-level normalization.

[1]  Faizan Javed,et al.  A pipeline for extracting and deduplicating domain-specific knowledge bases , 2015, 2015 IEEE International Conference on Big Data (Big Data).

[2]  Udo Hahn,et al.  High-performance gene name normalization with GENO , 2009, Bioinform..

[3]  Siddhartha Jonnalagadda,et al.  NEMO: Extraction and normalization of organization names from PubMed affiliation strings , 2010, Journal of biomedical discovery and collaboration.

[4]  Faizan Javed,et al.  An Ensemble Blocking Approach for Entity Resolution of Heterogeneous Datasets , 2017, FLAIRS Conference.

[5]  L. Jost Entropy and diversity , 2006 .

[6]  Jack Minker,et al.  An Analysis of Some Graph Theoretical Cluster Techniques , 1970, JACM.

[7]  Zhiyong Lu,et al.  DNorm: disease name normalization with pairwise learning to rank , 2013, Bioinform..

[8]  Tian Zhang,et al.  BIRCH: an efficient data clustering method for very large databases , 1996, SIGMOD '96.

[9]  Faizan Javed,et al.  Large-Scale Occupational Skills Normalization for Online Recruitment , 2017, AAAI.

[10]  Andrew Borthwick,et al.  Dynamic Record Blocking: Efficient Linking of Massive Databases in MapReduce , 2012 .

[11]  Erhard Rahm,et al.  Iterative Computation of Connected Graph Components with MapReduce , 2014, Datenbank-Spektrum.

[12]  Hans-Peter Kriegel,et al.  A Density-Based Algorithm for Discovering Clusters in Large Spatial Databases with Noise , 1996, KDD.

[13]  Malay K. Pakhira,et al.  A Linear Time-Complexity k-Means Algorithm Using Cluster Shifting , 2014, 2014 International Conference on Computational Intelligence and Communication Networks.

[14]  Jing Jiang,et al.  Linking Entities to a Knowledge Base with Query Expansion , 2011, EMNLP.

[15]  Yingjie Tian,et al.  A Comprehensive Survey of Clustering Algorithms , 2015, Annals of Data Science.

[16]  Jiawei Han,et al.  Entity Linking with a Knowledge Base: Issues, Techniques, and Solutions , 2015, IEEE Transactions on Knowledge and Data Engineering.

[17]  Mohammed J. Zaki Data Mining and Analysis: Fundamental Concepts and Algorithms , 2014 .

[18]  Siddhartha Jonnalagadda,et al.  NEMO: Extraction and normalization of organization names from PubMed affiliations , 2010, Journal of Biomedical Discovery and Collaboration.

[19]  Hae-Sang Park,et al.  A simple and fast algorithm for K-medoids clustering , 2009, Expert Syst. Appl..

[20]  Walid Magdy,et al.  Arabic Cross-Document Person Name Normalization , 2007, SEMITIC@ACL.

[21]  Anmol Bhasin,et al.  Entity Resolution Using Social Graphs for Business Applications , 2011, 2011 International Conference on Advances in Social Networks Analysis and Mining.

[22]  Praveen Paritosh,et al.  Freebase: a collaboratively created graph database for structuring human knowledge , 2008, SIGMOD Conference.

[23]  Ronnie Johansson,et al.  Choosing DBSCAN Parameters Automatically using Differential Evolution , 2014 .

[24]  Claudio Carpineto,et al.  A Survey of Automatic Query Expansion in Information Retrieval , 2012, CSUR.

[25]  W. Bruce Croft,et al.  Linear feature-based models for information retrieval , 2007, Information Retrieval.

[26]  Faizan Javed,et al.  CompanyDepot: Employer Name Normalization in the Online Recruitment Industry , 2016, KDD.

[27]  Anil K. Jain,et al.  Data clustering: a review , 1999, CSUR.

[28]  Faizan Javed,et al.  sCooL: A system for academic institution name normalization , 2014, 2014 International Conference on Collaboration Technologies and Systems (CTS).

[29]  Chih-Jen Lin,et al.  LIBSVM: A library for support vector machines , 2011, TIST.

[30]  Peter Christen,et al.  Data Matching , 2012, Data-Centric Systems and Applications.

[31]  Hakan Kardes,et al.  Graph-based Approaches for Organization Entity Resolution in MapReduce , 2013, TextGraphs@EMNLP.

[32]  Thomas Seidl,et al.  CC-MR - Finding Connected Components in Huge Graphs with MapReduce , 2012, ECML/PKDD.