Non-parametric topic model for discovering geographical topic variations

This paper presents a non-parametric topic model that captures not only the latent topics in text collections, but also how the topics change over space. Unlike other recent work that relies on either Gaussian assumptions or discretization of locations, here topics are associated with a distance dependent Chinese Restaurant Process (ddCRP), and for each document, the observed words are influenced by the document’s GPS-tag. Our model allows both unbound number and flexible distribution of the geographical variations of the topics’ content. We develop a Gibbs sampler for the proposal, and compare it with existing models on a real data set basis.