Font Size: a A A

Beyond bags of words: Effectively modeling dependence and features in information retrieval

Posted on:2008-08-14Degree:Ph.DType:Dissertation
University:University of Massachusetts AmherstCandidate:Metzler, Donald A., JrFull Text:PDF
GTID:1448390005956329Subject:Computer Science
Abstract/Summary:PDF Full Text Request
Current state of the art information retrieval models treat documents and queries as bags of words. There have been many attempts to go beyond this simple representation. Unfortunately, few have shown consistent improvements in retrieval effectiveness across a wide range of tasks and data sets. Here, we propose a new statistical model for information retrieval based on Markov random fields. The proposed model goes beyond the bag of words assumption by allowing dependencies between terms to be incorporated into the model. This allows for a variety of textual and non-textual features to be easily combined under the umbrella of a single model. Within this framework, we explore the theoretical issues involved, parameter estimation, feature selection, and query expansion. We give experimental results from a number of information retrieval tasks, such as ad hoc retrieval and web search.
Keywords/Search Tags:Information retrieval, Model, Words
PDF Full Text Request
Related items