论文信息 - Offline Comparison of Ranking Functions using Randomized Data

Offline Comparison of Ranking Functions using Randomized Data

Ranking functions return ranked lists of items, and users often interact with these items. How to evaluate ranking functions using historical interaction logs, also known as off-policy evaluation, is an important but challenging problem. The commonly used Inverse Propensity Scores (IPS) approaches work better for the single item case, but suffer from extremely low data efficiency for the ranked list case. In this paper, we study how to improve the data efficiency of IPS approaches in the offline comparison setting. We propose two approaches Trunc-match and Rand-interleaving for offline comparison using uniformly randomized data. We show that these methods can improve the data efficiency and also the comparison sensitivity based on one of the largest email search engines.

Michael Bendersky | Marc Najork | Aman Agarwal | Xuanhui Wang | Cheng Li

[1] Thorsten Joachims,et al. Unbiased Learning-to-Rank with Biased Feedback , 2016, WSDM.

[2] John Langford,et al. Off-policy evaluation for slate recommendation , 2016, NIPS.

[3] W. Bruce Croft,et al. Unbiased Learning to Rank with Unbiased Propensity Estimation , 2018, SIGIR.

[4] Thorsten Joachims,et al. Effective Evaluation Using Logged Bandit Feedback from Multiple Loggers , 2017, KDD.

[5] Thorsten Joachims,et al. Unbiased Comparative Evaluation of Ranking Functions , 2016, ICTIR.

[6] Lihong Li,et al. Toward Predicting the Outcome of an A/B Experiment for Search Relevance , 2015, WSDM.

[7] Wei Chu,et al. Unbiased offline evaluation of contextual-bandit-based news article recommendation algorithms , 2010, WSDM '11.

[8] Ben Carterette,et al. Offline Comparative Evaluation with Incremental, Minimally-Invasive Online Feedback , 2018, SIGIR.

[9] Marc Najork,et al. Learning to Rank with Selection Bias in Personal Search , 2016, SIGIR.

[10] Filip Radlinski,et al. Online Evaluation for Information Retrieval , 2016, Found. Trends Inf. Retr..

[11] Joaquin Quiñonero Candela,et al. Counterfactual reasoning and learning systems: the example of computational advertising , 2013, J. Mach. Learn. Res..

[12] Marc Najork,et al. Position Bias Estimation for Unbiased Learning to Rank in Personal Search , 2018, WSDM.

[13] Filip Radlinski,et al. How does clickthrough data reflect retrieval quality? , 2008, CIKM '08.