Title language model for information retrieval

11 August 2002

proceedings article
Published by Association for Computing Machinery (ACM)

p. 42-48
https://doi.org/10.1145/564376.564386

Abstract

In this paper, we propose a new language model, namely, a title language model, for information retrieval. Different from the traditional language model used for retrieval, we define the conditional probability P(Q|D) as the probability of using query Q as the title for document D. We adopted the statistical translation model learned from the title and document pairs in the collection to compute the probability P(Q|D). To avoid the sparse data problem, we propose two new smoothing methods. In the experiments with four different TREC document collections, the title language model for information retrieval with the new smoothing method outperforms both the traditional language model and the vector space model for IR significantly.

Keywords

This publication has 7 references indexed in Scilit:

A study of smoothing methods for language models applied to Ad Hoc information retrieval
Published by Association for Computing Machinery (ACM) ,2001
Relevance based language models
Published by Association for Computing Machinery (ACM) ,2001
Document language models, query models, and risk minimization for information retrieval
Published by Association for Computing Machinery (ACM) ,2001
Applying summarization techniques for term selection in relevance feedback
Published by Association for Computing Machinery (ACM) ,2001
Information retrieval as statistical translation
Published by Association for Computing Machinery (ACM) ,1999
A hidden Markov model information retrieval system
Published by Association for Computing Machinery (ACM) ,1999
A language modeling approach to information retrieval
Published by Association for Computing Machinery (ACM) ,1998