On the Estimation and Use of Statistical Modelling in Information Retrieval
arXiv:1904.00289
Abstract
Several tasks in information retrieval (IR) rely on assumptions regarding the distribution of some property (such as term frequency) in the data being processed. This thesis argues that such distributional assumptions can lead to incorrect conclusions and proposes a statistically principled method for determining the "true" distribution. This thesis further applies this method to derive a new family of ranking models that adapt their computations to the statistics of the data being processed. Experimental evaluation shows results on par or better than multiple strong baselines on several TREC collections. Overall, this thesis concludes that distributional assumptions can be replaced with an effective, efficient and principled method for determining the "true" distribution and that using the "true" distribution can lead to improved retrieval performance.
Phd thesis
References in corpus (9)
- Uncovering the overlapping community structure of complex networks in nature and society
- Finding community structure in networks using the eigenvectors of matrices
- Collaborative Tagging and Semiotic Dynamics
- Frequentist statistics as a theory of inductive inference
- Collaborative thesaurus tagging the Wikipedia way
- Where do statistical models come from? Revisiting the problem of specification
- Vocabulary growth in collaborative tagging systems
- On the proficient use of GEV distribution: a case study of subtropical monsoon region in India
- Power Law of Customers' Expenditures in Convenience Stores