Exploiting the hierarchical structure for link analysis
- 15 August 2005
- conference paper
- Published by Association for Computing Machinery (ACM)
- p. 186-193
- https://doi.org/10.1145/1076034.1076068
Abstract
Link analysis algorithms have been extensively used in Web information retrieval. However, current link analysis algorithms generally work on a flat link graph, ignoring the hierarchal structure of the Web graph. They often suffer from two problems: the sparsity of link graph and biased ranking of newly-emerging pages. In this paper, we propose a novel ranking algorithm called Hierarchical Rank as a solution to these two problems, which considers both the hierarchical structure and the link structure of the Web. In this algorithm, Web pages are first aggregated based on their hierarchical structure at directory, host or domain level and link analysis is performed on the aggregated graph. Then, the importance of each node on the aggregated graph is distributed to individual pages belong to the node based on the hierarchical structure. This algorithm allows the importance of linked Web pages to be distributed in the Web page space even when the space is sparse and contains new pages. Experimental results on the .GOV collection of TREC 2003 and 2004 show that hierarchical ranking algorithm consistently outperforms other well-known ranking algorithms, including the PageRank, BlockRank and LayerRank. In addition, experimental results show that link aggregation at the host level is much better than link aggregation at either the domain or directory levels.Keywords
This publication has 9 references indexed in Scilit:
- Block-level link analysisPublished by Association for Computing Machinery (ACM) ,2004
- Ranking the web frontierPublished by Association for Computing Machinery (ACM) ,2004
- Efficient pagerank approximation via graph aggregationPublished by Association for Computing Machinery (ACM) ,2004
- Topic-sensitive PageRankPublished by Association for Computing Machinery (ACM) ,2002
- Enhanced topic distillation using text, markup tags, and hyperlinksPublished by Association for Computing Machinery (ACM) ,2001
- Authoritative sources in a hyperlinked environmentJournal of the ACM, 1999
- On power-law relationships of the Internet topologyPublished by Association for Computing Machinery (ACM) ,1999
- Improved algorithms for topic distillation in a hyperlinked environmentPublished by Association for Computing Machinery (ACM) ,1998
- Overview of the Okapi projectsJournal of Documentation, 1997