Documents, such as those seen on Wikipedia and Folksonomy, have tended to be assigned with multiple topics as a meta-data.Therefore, it is more and more important to analyze a relationship between a document and topics assigned to the document. In this paper, we proposed a novel probabilistic generative model of documents with multiple topics as a meta-data. By focusing on modeling the generation process of a document with multiple topics, we can extract specific properties of documents with multiple topics.Proposed model is an expansion of an existing probabilistic generative model: Parametric Mixture Model (PMM). PMM models documents with multiple topics by mixing model parameters of each single topic. Since, however, PMM assigns the same mixture ratio to each single topic, PMM cannot take into account the bias of each topic within a document. To deal with this problem, we propose a model that considers Dirichlet distribution as a prior distribution of the mixture ratio.We adopt Variational Bayes Method to infer the bias of each topic within a document. We evaluate the proposed model and PMM using MEDLINE corpus.The results of F-measure, Precision and Recall show that the proposed model is more effective than PMM on multiple-topic classification. Moreover, we indicate the potential of the proposed model that extracts topics and document-specific keywords using information about the assigned topics.
[1]
Yiming Yang,et al.
A Comparative Study on Feature Selection in Text Categorization
,
1997,
ICML.
[2]
Nasser M. Nasrabadi,et al.
Pattern Recognition and Machine Learning
,
2006,
Technometrics.
[3]
Christopher M. Bishop,et al.
Pattern Recognition and Machine Learning (Information Science and Statistics)
,
2006
.
[4]
H. Attias,et al.
Learning parameters and structure of latent variable models by variation Bayes
,
1999
.
[5]
Michael I. Jordan,et al.
Hierarchical Dirichlet Processes
,
2006
.
[6]
Michael I. Jordan,et al.
Latent Dirichlet Allocation
,
2001,
J. Mach. Learn. Res..
[7]
Naonori Ueda,et al.
Single-shot detection of multiple categories of text using parametric mixture models
,
2002,
KDD.
[8]
Hinrich Schütze,et al.
Book Reviews: Foundations of Statistical Natural Language Processing
,
1999,
CL.
[9]
Hagai Attias,et al.
Inferring Parameters and Structure of Latent Variable Models by Variational Bayes
,
1999,
UAI.