论文信息 - From Predictive to Prescriptive Analytics

From Predictive to Prescriptive Analytics

In this paper, we combine ideas from machine learning (ML) and operations research and management science (OR/MS) in developing a framework, along with specific methods, for using data to prescribe optimal decisions in OR/MS problems. In a departure from other work on data-driven optimization and reflecting our practical experience with the data available in applications of OR/MS, we consider data consisting, not only of observations of quantities with direct effect on costs/revenues, such as demand or returns, but predominantly of observations of associated auxiliary quantities. The main problem of interest is a conditional stochastic optimization problem, given imperfect observations, where the joint probability distributions that specify the problem are unknown. We demonstrate that our proposed solution methods, which are inspired by ML methods such as local regression, CART, and random forests, are generally applicable to a wide range of decision problems. We prove that they are tractable and asymptotically optimal even when data is not iid and may be censored. We extend this to the case where decision variables may directly affect uncertainty in unknown ways, such as pricing's effect on demand. As an analogue to R^2, we develop a metric P termed the coefficient of prescriptiveness to measure the prescriptive content of data and the efficacy of a policy from an operations perspective. To demonstrate the power of our approach in a real-world setting we study an inventory management problem faced by the distribution arm of an international media conglomerate, which ships an average of 1bil units per year. We leverage internal data and public online data harvested from IMDb, Rotten Tomatoes, and Google to prescribe operational decisions that outperform baseline measures. Specifically, the data we collect, leveraged by our methods, accounts for an 88\% improvement as measured by our P.

Dimitris Bertsimas | Nathan Kallus | D. Bertsimas | Nathan Kallus

[1] E. L. Lehmann,et al. Theory of point estimation , 1950 .

[2] Vladimir Vapnik,et al. Principles of Risk Minimization for Learning Theory , 1991, NIPS.

[3] E. Kaplan,et al. Nonparametric Estimation from Incomplete Observations , 1958 .

[4] Eleftherios Mylonakis,et al. Google trends: a web-based tool for real-time surveillance of disease outbreaks. , 2009, Clinical infectious diseases : an official publication of the Infectious Diseases Society of America.

[5] E. Nadaraya. On Estimating Regression , 1964 .

[6] P. Doukhan. Mixing: Properties and Examples , 1994 .

[7] Dimitri P. Bertsekas,et al. Dynamic Programming and Optimal Control, Two Volume Set , 1995 .

[8] Dimitri P. Bertsekas,et al. Nonlinear Programming , 1997 .

[9] M. Talagrand,et al. Probability in Banach Spaces: Isoperimetry and Processes , 1991 .

[10] Jianqing Fan. Local Linear Regression Smoothers and Their Minimax Efficiencies , 1993 .

[11] P. Billingsley,et al. Convergence of Probability Measures , 1969 .

[12] R. Tibshirani. Regression Shrinkage and Selection via the Lasso , 1996 .

[13] Ambuj Tewari,et al. On the Complexity of Linear Prediction: Risk Bounds, Margin Bounds, and Regularization , 2008, NIPS.

[14] G. Calafiore,et al. On Distributionally Robust Chance-Constrained Linear Programs , 2006 .

[15] Wei-Yin Loh,et al. Classification and regression trees , 2011, WIREs Data Mining Knowl. Discov..

[16] Xiaohong Chen,et al. MIXING AND MOMENT PROPERTIES OF VARIOUS GARCH AND STOCHASTIC VOLATILITY MODELS , 2002, Econometric Theory.

[17] George G. Roussas,et al. Exact rates of almost sure convergence of a recursive kernel estimate of a probability densiy function: Application to regression and hazard rate estimation , 1992 .

[18] Pierre Geurts,et al. Extremely randomized trees , 2006, Machine Learning.

[19] R. C. Bradley. Basic Properties of Strong Mixing Conditions , 1985 .

[20] Alexander Shapiro,et al. The Sample Average Approximation Method for Stochastic Discrete Optimization , 2002, SIAM J. Optim..

[21] A. Mokkadem. Mixing properties of ARMA processes , 1988 .

[22] Cosma Rohilla Shalizi,et al. Estimating beta-mixing coefficients , 2011, AISTATS.

[23] B. Hansen. UNIFORM CONVERGENCE RATES FOR KERNEL ESTIMATION WITH DEPENDENT DATA , 2008, Econometric Theory.

[24] Bernardo A. Huberman,et al. Predicting the Future with Social Media , 2010, 2010 IEEE/WIC/ACM International Conference on Web Intelligence and Intelligent Agent Technology.

[25] Leo Breiman,et al. Random Forests , 2001, Machine Learning.

[26] Shuaian Wang,et al. Sample Average Approximation , 2018 .

[27] E. F. Schuster,et al. On a universal strong law of large numbers for conditional expectations , 1998 .

[28] Don R. Hush,et al. An Explicit Description of the Reproducing Kernel Hilbert Spaces of Gaussian RBF Kernels , 2006, IEEE Transactions on Information Theory.

[29] David M. Pennock,et al. Predicting consumer behavior with Web search , 2010, Proceedings of the National Academy of Sciences.

[30] Yinyu Ye,et al. Distributionally Robust Optimization Under Moment Uncertainty with Application to Data-Driven Problems , 2010, Oper. Res..

[31] H. Robbins. Some aspects of the sequential design of experiments , 1952 .

[32] Leo Breiman,et al. Statistical Modeling: The Two Cultures (with comments and a rejoinder by the author) , 2001 .

[33] Jean-Philippe Vial,et al. Robust Optimization , 2021, ICORES.

[34] Antonio Alonso Ayuso,et al. Introduction to Stochastic Programming , 2009 .

[35] Erik Brynjolfsson,et al. Big data: the management revolution. , 2012, Harvard business review.

[36] YeYinyu,et al. Distributionally Robust Optimization Under Moment Uncertainty with Application to Data-Driven Problems , 2010 .