On the design of optimal computer experiments to model solvent effects on reaction kinetics

被引:1
|
作者
Gui, Lingfeng [1 ,2 ]
Armstrong, Alan [3 ,4 ]
Galindo, Amparo [1 ,2 ]
Sayyed, Fareed Bhasha [5 ]
Kolis, Stanley P. [6 ]
Adjiman, Claire S. [1 ,2 ]
机构
[1] Imperial Coll London, Sargent Ctr Proc Syst Engn, Dept Chem Engn, London SW7 2AZ, England
[2] Imperial Coll London, Inst Mol Sci & Engn, London SW7 2AZ, England
[3] Imperial Coll London, Dept Chem, White City Campus, London W12 0BZ, England
[4] Imperial Coll London, Inst Mol Sci & Engn, Mol Sci Res Hub, White City Campus, London W12 0BZ, England
[5] Eli Lilly Serv India Pvt Ltd, Synthet Mol Design & Dev, Bengaluru 560103, India
[6] Eli Lilly & Co, Lilly Corp Ctr, Synthet Mol Design & Dev, Indianapolis, IN 46285 USA
来源
基金
英国工程与自然科学研究理事会;
关键词
SOLVATION ENERGY RELATIONSHIPS; SOLVATOCHROMIC PARAMETERS; CONSTANTS; SELECTION;
D O I
10.1039/d4me00074a
中图分类号
O64 [物理化学(理论化学)、化学物理学];
学科分类号
070304 ; 081704 ;
摘要
Developing an accurate predictive model of solvent effects on reaction kinetics is a challenging task, yet it can play an important role in process development. While first-principles or machine learning models are often compute- or data-intensive, simple surrogate models, such as multivariate linear or quadratic regression models, are useful when computational resources and data are scarce. The judicious choice of a small set of training data, i.e., a set of solvents in which quantum mechanical (QM) calculations of liquid-phase rate constants are to be performed, is critical to obtaining a reliable model. This is, however, made especially challenging by the highly irregular shape of the discrete space of possible experiments (solvent choices). In this work, we demonstrate that when choosing a set of computer experiments to generate training data, the D-optimality criterion value of the chosen set correlates well with the likelihood of achieving good model performance. With the Menshutkin reaction of pyridine and phenacyl bromide as a case study, this finding is further verified via the evaluation of the surrogate models regressed using D-optimal solvent sets generated from four distinct selection spaces. We also find that incorporating quadratic terms in the surrogate model and choosing the D-optimal solvent set from a selection space similar to the test set can significantly improve the accuracy of reaction rate constant predictions while using a small training dataset. Our approach holds promise for the use of statistical optimality criteria for other types of computer experiments, supporting the construction of surrogate models with reduced resource and data requirements. Model-based design of experiments using the D-optimality criterion can help select computer experiments to generate more information-rich training sets and leads to more reliable surrogate models that can be used for efficient molecular design.
引用
收藏
页码:1254 / 1274
页数:21
相关论文
共 50 条
  • [22] D-optimal design of DSC experiments for nth order kinetics
    Aravind Manerswammy
    Stuart H. Munson-McGee
    Robert Steiner
    Charles L. Johnson
    Journal of Thermal Analysis and Calorimetry, 2009, 97 : 895 - 902
  • [23] D-optimal design of DSC experiments for nth order kinetics
    Manerswammy, Aravind
    Munson-McGee, Stuart H.
    Steiner, Robert
    Johnson, Charles L.
    JOURNAL OF THERMAL ANALYSIS AND CALORIMETRY, 2009, 97 (03) : 895 - 902
  • [24] Computer-aided reaction solvent design considering inertness using group contribution-based reaction thermodynamic model
    Liu, Qilei
    Zhang, Lei
    Tang, Kun
    Feng, Yixuan
    Zhang, Jinyuan
    Zhuang, Yu
    Liu, Linlin
    Du, Jian
    CHEMICAL ENGINEERING RESEARCH & DESIGN, 2019, 152 : 123 - 133
  • [25] A QM-CAMD approach to solvent design for optimal reaction rates
    Struebing, Heiko
    Obermeier, Stephan
    Siougkrou, Eirini
    Adjiman, Claire S.
    Galindo, Amparo
    CHEMICAL ENGINEERING SCIENCE, 2017, 159 : 69 - 83
  • [26] A computer-aided methodology for optimal solvent design for reactions with experimental verification
    Folic, M
    Adjiman, CS
    Pistikopoulos, EN
    European Symposium on Computer-Aided Process Engineering-15, 20A and 20B, 2005, 20a-20b : 1651 - 1656
  • [27] COMPUTER-AIDED MOLECULAR DESIGN - A NOVEL METHOD FOR OPTIMAL SOLVENT SELECTION
    ODELE, O
    MACCHIETTO, S
    FLUID PHASE EQUILIBRIA, 1993, 82 : 47 - 54
  • [28] Optimal design of stimulus experiments for robust discrimination of biochemical reaction networks
    Flassig, R. J.
    Sundmacher, K.
    BIOINFORMATICS, 2012, 28 (23) : 3089 - 3096
  • [29] Sequential design for computer experiments with a flexible Bayesian additive model
    Chipman, Hugh
    Ranjan, Pritam
    Wang, Weiwei
    CANADIAN JOURNAL OF STATISTICS-REVUE CANADIENNE DE STATISTIQUE, 2012, 40 (04): : 663 - 678
  • [30] Sequential Design of Computer Experiments for the Computation of Bayesian Model Evidence
    Sinsbeck, Michael
    Cooke, Emily
    Nowak, Wolfgang
    SIAM-ASA JOURNAL ON UNCERTAINTY QUANTIFICATION, 2021, 9 (01): : 260 - 279