The topic information of conversational content is important for continuation with communication, so topic detection and tracking is one of important research. Due to there are many topic transform occurring frequently in long time communication, and the conversation maybe have many topics, so it's important to detect different topics in conversational content. This paper detects topic information by using agglomerative clustering of utterances and Dynamic Latent Dirichlet Allocation topic model, uses proportion of verb and noun to analyze similarity between utterances and cluster all utterances in conversational content by agglomerative clustering algorithm. The topic structure of conversational content is friability, so we use speech act information and gets the hypernym information by E-HowNet that obtains robustness of word categories. Latent Dirichlet Allocation topic model is used to detect topic in file units, it just can detect only one topic if uses it in conversational content, because of there are many topics in conversational content frequently, and also uses speech act information and hypernym information to train the latent Dirichlet allocation models, then uses trained models to detect different topic information in conversational content. For evaluating the proposed method, support vector machine is developed for comparison. According to the experimental results, we can find the proposed method outperforms the approach based on support vector machine in topic detection and tracking in spoken dialogue.
[1]
Chih-Jen Lin,et al.
LIBSVM: A library for support vector machines
,
2011,
TIST.
[2]
Michael I. Jordan,et al.
Latent Dirichlet Allocation
,
2001,
J. Mach. Learn. Res..
[3]
Yaw-Huei Chen,et al.
Extracting Topics Information from Conference Web Pages Using Page Segmentation and SVM
,
2010,
2010 International Conference on Technologies and Applications of Artificial Intelligence.
[4]
Corinna Cortes,et al.
Support-Vector Networks
,
1995,
Machine Learning.
[5]
Deyuan Zhang,et al.
Finding main topics in blogosphere using document clustering based on topic model
,
2011,
2011 International Conference on Machine Learning and Cybernetics.
[6]
Xiaolong Wang,et al.
A comparative study of topic models for topic clustering of Chinese web news
,
2010,
2010 3rd International Conference on Computer Science and Information Technology.
[7]
Yan Chen,et al.
A topic detection method based on Semantic Dependency Distance and PLSA
,
2012,
Proceedings of the 2012 IEEE 16th International Conference on Computer Supported Cooperative Work in Design (CSCWD).