High-Efficiency Machine Learning Method for Identifying Foodborne Disease Outbreaks and Confounding Factors

被引:12
|
作者
Zhang, Peng [1 ,2 ]
Cui, Wenjuan [1 ]
Wang, Hanxue [1 ,2 ]
Du, Yi [1 ,2 ]
Zhou, Yuanchun [1 ,2 ]
机构
[1] Chinese Acad Sci, Comp Network Informat Ctr, Bldg 2,Software Pk 4,South Fourth St, Beijing 100190, Peoples R China
[2] Univ Chinese Acad Sci, Sch Comp Sci & Technol, Beijing, Peoples R China
关键词
foodborne disease outbreaks; machine learning; foodborne disease; SURVEILLANCE;
D O I
10.1089/fpd.2020.2913
中图分类号
TS2 [食品工业];
学科分类号
0832 ;
摘要
The China National Center for Food Safety Risk Assessment (CFSA) uses the Foodborne Disease Monitoring and Reporting System (FDMRS) to monitor outbreaks of foodborne diseases across the country. However, there are problems of underreporting or erroneous reporting in FDMRS, which significantly increase the cost of related epidemic investigations. To solve this problem, we designed a model to identify suspected outbreaks from the data generated by the FDMRS of CFSA. In this study, machine learning models were used to fit the data. The recall rate and F1-score were used as evaluation metrics to compare the classification performance of each model. Feature importance and pathogenic factors were identified and analyzed using tree-based and gradient boosting models. Three real foodborne disease outbreaks were then used to evaluate the best performing model. Furthermore, the SHapley Additive exPlanation value was used to identify the effect of features. Among all machine learning classification models, the eXtreme Gradient Boosting (XGBoost) model achieved the best performance, with the highest recall rate and F1-score of 0.9699 and 0.9582, respectively. In terms of model validation, the model provides a correct judgment of real outbreaks. In the feature importance analysis with the XGBoost model, the health status of the other people with the same exposure has the highest weight, reaching 0.65. The machine learning model built in this study exhibits high accuracy in recognizing foodborne disease outbreaks, thus reducing the manual burden for medical staff. The model helped us identify the confounding factors of foodborne disease outbreaks. Attention should be paid not only to the health status of those with the same exposure but also to the similarity of the cases in time and space.
引用
收藏
页码:590 / 598
页数:9
相关论文
共 50 条
  • [1] FACTORS THAT CONTRIBUTE TO OUTBREAKS OF FOODBORNE DISEASE
    BRYAN, FL
    [J]. JOURNAL OF FOOD PROTECTION, 1978, 41 (10) : 816 - 827
  • [2] Identifying Pathogens of Foodborne Diseases with Machine Learning
    Wang, Hanxue
    Cui, Wenjuan
    Zhou, Yuanchun
    Du, Yi
    [J]. Data Analysis and Knowledge Discovery, 2021, 5 (09) : 54 - 62
  • [3] Predicting Foodborne Disease Outbreaks with Food Safety Certifications: Econometric and Machine Learning Analyses
    Zheng, Yuqing
    Gracia, Azucena
    Hu, Lijiao
    [J]. JOURNAL OF FOOD PROTECTION, 2023, 86 (09)
  • [4] Machine learning helps identifying relations and confounding factors in radiomics-based models
    Traverso, A.
    Kazmierski, M.
    Wee, L.
    Dekker, A.
    Welch, M.
    Hosni, A.
    Jaffray, D.
    Hope, A.
    [J]. RADIOTHERAPY AND ONCOLOGY, 2019, 133 : S162 - S163
  • [5] Prospective Detection of Foodborne Illness Outbreaks Using Machine Learning Approaches
    Teyhouee, Aydin
    McPhee-Knowles, Sara
    Waldner, Chryl
    Osgood, Nathaniel
    [J]. SOCIAL, CULTURAL, AND BEHAVIORAL MODELING, 2017, 10354 : 302 - 308
  • [6] Unlocking HDR-mediated nucleotide editing by identifying high-efficiency target sites using machine learning
    O'Brien, Aidan R.
    Wilson, Laurence O. W.
    Burgio, Gaetan
    Bauer, Denis C.
    [J]. SCIENTIFIC REPORTS, 2019, 9 (1)
  • [7] Unlocking HDR-mediated nucleotide editing by identifying high-efficiency target sites using machine learning
    Aidan R. O’Brien
    Laurence O. W. Wilson
    Gaetan Burgio
    Denis C. Bauer
    [J]. Scientific Reports, 9
  • [8] Identifying factors associated with periodontal disease using machine learning
    Alqahtani, Hussam M.
    Koroukian, Siran M.
    Stange, Kurt
    Schiltz, Nicholas K.
    Bissada, Nabil F.
    [J]. JOURNAL OF INTERNATIONAL SOCIETY OF PREVENTIVE AND COMMUNITY DENTISTRY, 2022, 12 (06): : 612 - 620
  • [9] Interpretable AI and Machine Learning Classification for Identifying High-Efficiency Donor-Acceptor Pairs in Organic Solar Cells
    Siddiqui, Hamza
    Usmani, Tahsin
    [J]. ACS OMEGA, 2024, 9 (32): : 34445 - 34455
  • [10] High-efficiency synthesis of red carbon dots using machine learning
    Luo, Jun Bo
    Chen, Jiao
    Liu, Hui
    Huang, Cheng Zhi
    Zhou, Jun
    [J]. CHEMICAL COMMUNICATIONS, 2022, 58 (64) : 9014 - 9017