Bayesian Meta-Prior Learning Using Empirical Bayes

被引:6
|
作者
Nabi, Sareh [1 ,2 ]
Nassif, Houssam [2 ]
Hong, Joseph [2 ]
Mamani, Hamed [1 ]
Imbens, Guido [2 ,3 ]
机构
[1] Univ Washington, Foster Sch Business, Seattle, WA 98195 USA
[2] Amazon, Seattle, WA 98109 USA
[3] Stanford Univ, Grad Sch Business, Stanford, CA 94305 USA
关键词
informative prior; meta-prior; empirical Bayes; Bayesian bandit; generalized linear models; Thompson sampling; feature grouping; learning rate;
D O I
10.1287/mnsc.2021.4136
中图分类号
C93 [管理学];
学科分类号
12 ; 1201 ; 1202 ; 120202 ;
摘要
Adding domain knowledge to a learning system is known to improve results. In multiparameter Bayesian frameworks, such knowledge is incorporated as a prior. On the other hand, the various model parameters can have different learning rates in real-world problems, especially with skewed data. Two often-faced challenges in operation management and management science applications are the absence of informative priors and the inability to control parameter learning rates. In this study, we propose a hierarchical empirical Bayes approach that addresses both challenges and that can generalize to any Bayesian framework. Our method learns empirical meta-priors from the data itself and uses them to decouple the learning rates of first-order and second-order features (or any other given feature grouping) in a generalized linear model. Because the first-order features are likely to have a more pronounced effect on the outcome, focusing on learning first-order weights first is likely to improve performance and convergence time. Our empirical Bayes method clamps features in each group together and uses the deployed model's observed data to empirically compute a hierarchical prior in hindsight. We report theoretical results for the unbiasedness, strong consistency, and optimal frequentist cumulative regret properties of our meta-prior variance estimator. We apply our method to a standard supervised learning optimization problem as well as an online combinatorial optimization problem in a contextual bandit setting implemented in an Amazon production system. During both simulations and live experiments, our method shows marked improvements, especially in cases of small traffic. Our findings are promising because optimizing over sparse data is often a challenge.
引用
收藏
页码:1737 / 1755
页数:20
相关论文
共 50 条
  • [1] EMPIRICAL BAYES WITH A CHANGING PRIOR
    MARA, MK
    DEELY, JJ
    ANNALS OF STATISTICS, 1984, 12 (03): : 1071 - 1078
  • [2] Adaptively leveraging external data with robust meta-analytical-predictive prior using empirical Bayes
    Zhang, Hongtao
    Shen, Yueqi
    Li, Judy
    Ye, Han
    Chiang, Alan Y. Y.
    PHARMACEUTICAL STATISTICS, 2023, 22 (05) : 846 - 860
  • [3] Bayes and empirical Bayes semi-blind deconvolution using eigenfunctions of a prior covariance
    Pillonetto, Gianluigi
    Bell, Bradley M.
    AUTOMATICA, 2007, 43 (10) : 1698 - 1712
  • [4] Bayesian Prior Choice in IRT Estimation Using MCMC and Variational Bayes
    Natesan, Prathiba
    Nandakumar, Ratna
    Minka, Tom
    Rubright, Jonathan D.
    FRONTIERS IN PSYCHOLOGY, 2016, 7
  • [5] Empirical Bayesian Inference Using a Support Informed Prior
    Zhang, Jiahui
    Gelb, Anne
    Scarnati, Theresa
    SIAM-ASA JOURNAL ON UNCERTAINTY QUANTIFICATION, 2022, 10 (02): : 745 - 774
  • [6] EMPIRICAL BAYES APPROACH - ESTIMATING PRIOR DISTRIBUTION
    RUTHERFORD, JR
    KRUTCHKOFF, RG
    BIOMETRIKA, 1967, 54 : 326 - +
  • [7] EMPIRICAL BAYES ESTIMATION OF PRIOR AND POSTERIOR DISTRIBUTION
    RUTHERFO.JR
    KRUTCHKO.RG
    ANNALS OF MATHEMATICAL STATISTICS, 1965, 36 (05): : 1606 - &
  • [8] Empirical Bayes methods in classical and Bayesian inference
    Petrone S.
    Rizzelli S.
    Rousseau J.
    Scricciolo C.
    METRON, 2014, 72 (2) : 201 - 215
  • [9] EMPIRICAL BAYES ESTIMATES USING THE NONPARAMETRIC MAXIMUM-LIKELIHOOD ESTIMATE FOR THE PRIOR
    LAIRD, NM
    JOURNAL OF STATISTICAL COMPUTATION AND SIMULATION, 1982, 15 (2-3) : 211 - 220
  • [10] Bayesian change-point problem using Bayes factor with hierarchical prior distribution
    Jung, Myoungjin
    Song, Seongho
    Chung, Younshik
    COMMUNICATIONS IN STATISTICS-THEORY AND METHODS, 2017, 46 (03) : 1352 - 1366