论文信息 - Prior-Knowledge-Embedded LDA with Word2vec - for Detecting Specific Topics in Documents

Prior-Knowledge-Embedded LDA with Word2vec - for Detecting Specific Topics in Documents

This paper proposes a method to apply prior knowledge about topics of interest to Latent Dirichlet Allocation (LDA). The conventional LDA sometimes fails to detect specific topics of interest. Therefore, our approach uses word2vec to acquire linkages between words related to specific topics. The extracted linkages are used as prior knowledge about the topics in the subsequent LDA process. The extracted linkages can also be used to annotate words in a consistent manner. Such consistent annotations cannot be realized using conventional LDA, which relies on bag-of-words–based clustering. We examine our approach by applying it to travelers’ reviews, to detect topics related to Japanese shrines. The experimental results show that our approach is effective in the following three aspects: (1) The average coherence of our approach, i.e., the semantic consistencies among words, outperforms that of the conventional LDA. (2) Words in each sentence are annotated such that the annotations reflect the topic of the sentence. The conventional LDA sometimes makes confusing/mixed annotations to the words in a single sentence. Our approach, on the contrary, can make annotations that reflect the topic of the sentence in a consistent manner. (3) Our approach enables to detect very specific topics complying with users’ interests.

Akihiro Ito | Kenichi Yoshida | Yutaka Saito | Hiroshi Uehara

[1] Michael I. Jordan,et al. Latent Dirichlet Allocation , 2001, J. Mach. Learn. Res..

[2] Yao Lu,et al. LDA Meets Word2Vec: A Novel Model for Academic Abstract Clustering , 2018, WWW.

[3] Yulan He,et al. Extracting Topical Phrases from Clinical Documents , 2016, AAAI.

[4] Rui Zhang,et al. Incorporating Knowledge Graph Embeddings into Topic Modeling , 2017, AAAI.

[5] Jeffrey Dean,et al. Distributed Representations of Words and Phrases and their Compositionality , 2013, NIPS.

[6] Frank Rudzicz,et al. Augmenting word2vec with latent Dirichlet allocation within a clinical application , 2019, NAACL-HLT.