Conditional Synthetic Data Generation for Robust Machine Learning Applications with Limited Pandemic Data

被引:0
|
作者
Das, Hari Prasanna [1 ]
Tran, Ryan [1 ]
Singh, Japjot [1 ]
Yue, Xiangyu [1 ]
Tison, Geoffrey [2 ]
Sangiovanni-Vincentelli, Alberto [1 ]
Spanos, Costas J. [1 ]
机构
[1] Univ Calif Berkeley, Dept Elect Engn & Comp Sci, Berkeley, CA 94720 USA
[2] Univ Calif San Francisco UCSF, Div Cardiol, San Francisco, CA USA
基金
新加坡国家研究基金会;
关键词
D O I
暂无
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Background: At the onset of a pandemic, such as COVID-19, data with proper labeling/attributes corresponding to the new disease might be unavailable or sparse. Machine Learning (ML) models trained with the available data, which is limited in quantity and poor in diversity, will often be biased and inaccurate. At the same time, ML algorithms designed to fight pandemics must have good performance and be developed in a time-sensitive manner. To tackle the challenges of limited data, and label scarcity in the available data, we propose generating conditional synthetic data, to be used alongside real data for developing robust ML models. Methods: We present a hybrid model consisting of a conditional generative flow and a classifier for conditional synthetic data generation. The classifier decouples the feature representation for the condition, which is fed to the flow to extract the local noise. We generate synthetic data by manipulating the local noise with fixed conditional feature representation. We also propose a semi-supervised approach to generate synthetic samples in the absence of labels for a majority of the available data. Results: We performed conditional synthetic generation for chest computed tomography (CT) scans corresponding to normal, COVID-19, and pneumonia afflicted patients. We show that our method significantly outperforms existing models both on qualitative and quantitative performance, and our semi-supervised approach can efficiently synthesize conditional samples under label scarcity. As an example of downstream use of synthetic data, we show improvement in COVID-19 detection from CT scans with conditional synthetic data augmentation.
引用
收藏
页码:11792 / 11800
页数:9
相关论文
共 50 条
  • [41] Data Mining and Machine Learning Applications for Educational Big Data in the University
    Abe, Keisuke
    IEEE 17TH INT CONF ON DEPENDABLE, AUTONOM AND SECURE COMP / IEEE 17TH INT CONF ON PERVAS INTELLIGENCE AND COMP / IEEE 5TH INT CONF ON CLOUD AND BIG DATA COMP / IEEE 4TH CYBER SCIENCE AND TECHNOLOGY CONGRESS (DASC/PICOM/CBDCOM/CYBERSCITECH), 2019, : 350 - 355
  • [42] Generation of Synthetic CPTs with Access to Limited Geotechnical Data for Offshore Sites
    Shoukat, Gohar
    Michel, Guillaume
    Coughlan, Mark
    Malekjafarian, Abdollah
    Thusyanthan, Indrasenan
    Desmond, Cian
    Pakrashi, Vikram
    ENERGIES, 2023, 16 (09)
  • [43] Augmented machine learning for sewage quality assessment with limited data
    Lv, Jia-Qiang
    Yin, Wan-Xin
    Xu, Jia-Min
    Cheng, Hao-Yi
    Li, Zhi-Ling
    Yang, Ji-Xian
    Wang, Ai-Jie
    Wang, Hong-Cheng
    ENVIRONMENTAL SCIENCE AND ECOTECHNOLOGY, 2025, 23
  • [44] Machine Learning for Large Scale Manufacturing Data with Limited Information
    Nedelkoski, S.
    Stojanovski, G.
    2017 13TH IEEE INTERNATIONAL CONFERENCE ON CONTROL & AUTOMATION (ICCA), 2017, : 70 - 75
  • [45] Machine learning for reactor power monitoring with limited labeled data
    Stewart, C. L.
    Goldblum, B. L.
    Abbott, R. G.
    Appleby, L.
    Borghetti, B. J.
    Hollingshead, V.
    Whetzel, J. H.
    NUCLEAR INSTRUMENTS & METHODS IN PHYSICS RESEARCH SECTION A-ACCELERATORS SPECTROMETERS DETECTORS AND ASSOCIATED EQUIPMENT, 2025, 1073
  • [46] Robust data-driven machine-learning models for subsurface applications are we there yet?
    Mishra, Srikanta
    Schuetter, Jared
    Datta-Gupta, Akhil
    Bromhal, Grant
    JPT, Journal of Petroleum Technology, 2021, 73 (03): : 25 - 30
  • [47] Machine Learning Models for Regional Photovoltaic Power Generation Forecasting with Limited Plant-Specific Data
    Tucci, Mauro
    Piazzi, Antonio
    Thomopulos, Dimitri
    ENERGIES, 2024, 17 (10)
  • [48] A systematic review and evaluation of synthetic simulated data generation strategies for deep learning applications in construction
    Xu, Liqun
    Liu, Hexu
    Xiao, Bo
    Luo, Xiaowei
    DharmarajVeeramani
    Zhu, Zhenhua
    ADVANCED ENGINEERING INFORMATICS, 2024, 62
  • [49] Machine learning and data mining: strategies for hypothesis generation
    Oquendo, M. A.
    Baca-Garcia, E.
    Artes-Rodriguez, A.
    Perez-Cruz, F.
    Galfalvy, H. C.
    Blasco-Fontecilla, H.
    Madigan, D.
    Duan, N.
    MOLECULAR PSYCHIATRY, 2012, 17 (10) : 956 - 959
  • [50] Machine learning and data mining: strategies for hypothesis generation
    M A Oquendo
    E Baca-Garcia
    A Artés-Rodríguez
    F Perez-Cruz
    H C Galfalvy
    H Blasco-Fontecilla
    D Madigan
    N Duan
    Molecular Psychiatry, 2012, 17 : 956 - 959