Learning to cluster web search results

Top Cited Papers

25 July 2004

proceedings article
Published by Association for Computing Machinery (ACM)

p. 210-217
https://doi.org/10.1145/1008992.1009030

Abstract

Organizing Web search results into clusters facilitates users' quick browsing through search results. Traditional clustering techniques are inadequate since they don't generate clusters with highly readable names. In this paper, we reformalize the clustering problem as a salient phrase ranking problem. Given a query and the ranked list of documents (typically a list of titles and snippets) returned by a certain Web search engine, our method first extracts and ranks salient phrases as candidate cluster names, based on a regression model learned from human labeled training data. The documents are assigned to relevant salient phrases to form candidate clusters, and the final clusters are generated by merging these candidate clusters. Experimental results verify our method's feasibility and effectiveness.

Keywords

This publication has 7 references indexed in Scilit:

Mining topic-specific concepts and definitions on the web
Published by Association for Computing Machinery (ACM) ,2003
Finding topic words for hierarchical summarization
Published by Association for Computing Machinery (ACM) ,2001
Grouper: a dynamic clustering interface to Web search results
Computer Networks, 1999
Web document clustering
Published by Association for Computing Machinery (ACM) ,1998
PAT-tree-based keyword extraction for Chinese information retrieval
Published by Association for Computing Machinery (ACM) ,1997
Reexamining the cluster hypothesis
Published by Association for Computing Machinery (ACM) ,1996
Constant interaction-time scatter/gather browsing of very large document collections
Published by Association for Computing Machinery (ACM) ,1993