Mining correlated bursty topic patterns from coordinated text streams
Top Cited Papers
- 12 August 2007
- proceedings article
- Published by Association for Computing Machinery (ACM)
- p. 784-793
- https://doi.org/10.1145/1281192.1281276
Abstract
Previous work on text mining has almost exclusively focused on a single stream. However, we often have available multiple text streams indexed by the same set of time points (called coordinated text streams), which offer new opportunities for text mining. For example, when a major event happens, all the news articles published by different agencies in different languages tend to cover the same event for a certain period, exhibiting a correlated bursty topic pattern in all the news article streams. In general, mining correlated bursty topic patterns from coordinated text streams can reveal interesting latent associations or events behind these streams. In this paper, we define and study this novel text mining problem. We propose a general probabilistic algorithm which can effectively discover correlated bursty patterns and their bursty periods across text streams even if the streams have completely different vocabularies (e.g., English vs Chinese). Evaluation of the proposed method on a news data set and a literature data set shows that it can effectively discover quite meaningful topic patterns from both data sets: the patterns discovered from the news data set accurately reveal the major common events covered in the two streams of news articles (in English and Chinese, respectively), while the patterns discovered from two database publication streams match well with the major research paradigm shifts in database research. Since the proposed method is general and does not require the streams to share vocabulary, it can be applied to any coordinated text streams to discover correlated topic patterns that burst in multiple streams in the same period.Keywords
This publication has 20 references indexed in Scilit:
- A mixture model for contextual text miningPublished by Association for Computing Machinery (ACM) ,2006
- Named entity transliteration with comparable corporaPublished by Association for Computational Linguistics (ACL) ,2006
- Mining comparable bilingual text corpora for cross-language information integrationPublished by Association for Computing Machinery (ACM) ,2005
- Semantic similarity between search engine queries using temporal correlationPublished by Association for Computing Machinery (ACM) ,2005
- A cross-collection mixture model for comparative text miningPublished by Association for Computing Machinery (ACM) ,2004
- On demand classification of data streamsPublished by Association for Computing Machinery (ACM) ,2004
- On the bursty evolution of blogspacePublished by Association for Computing Machinery (ACM) ,2003
- Bursty and hierarchical structure in streamsPublished by Association for Computing Machinery (ACM) ,2002
- Improving text categorization methods for event trackingPublished by Association for Computing Machinery (ACM) ,2000
- Extracting significant time varying features from textPublished by Association for Computing Machinery (ACM) ,1999