Scalable de novo classification of antibiotic resistance of Mycobacterium tuberculosis

被引:0
|
作者
Serajian, Mohammadali [1 ]
Marini, Simone [2 ]
Alanko, Jarno N. [3 ]
Noyes, Noelle R. [4 ]
Prosperi, Mattia [2 ]
Boucher, Christina [1 ]
机构
[1] Univ Florida, Dept Comp & Informat Sci & Engn, 1889 Museum Rd, Gainesville, FL 32611 USA
[2] Univ Florida, Dept Epidemiol, POB 100231, Gainesville, FL 32601 USA
[3] Univ Helsinki, Dept Comp Sci, POB 4, Helsinki 00014, Finland
[4] Univ Minnesota, Dept Vet Populat Med, 1365 Gortner Ave, St Paul, MN 55108 USA
关键词
READ ALIGNMENT; GENOME; TOOL;
D O I
10.1093/bioinformatics/btae243
中图分类号
Q5 [生物化学];
学科分类号
071010 ; 081704 ;
摘要
Motivation: World Health Organization estimates that there were over 10 million cases of tuberculosis (TB) worldwide in 2019, resulting in over 1.4 million deaths, with a worrisome increasing trend yearly. The disease is caused by Mycobacterium tuberculosis (MTB) through airborne transmission. Treatment of TB is estimated to be 85% successful, however, this drops to 57% if MTB exhibits multiple antimicrobial resistance (AMR), for which fewer treatment options are available. Results: We develop a robust machine-learning classifier using both linear and nonlinear models (i.e. LASSO logistic regression (LR) and random forests (RF)) to predict the phenotypic resistance of Mycobacterium tuberculosis (MTB) for a broad range of antibiotic drugs. We use data from the CRyPTIC consortium to train our classifier, which consists of whole genome sequencing and antibiotic susceptibility testing (AST) phenotypic data for 13 different antibiotics. To train our model, we assemble the sequence data into genomic contigs, identify all unique 31-mers in the set of contigs, and build a feature matrix M, where M[i, j] is equal to the number of times the ith 31-mer occurs in the jth genome. Due to the size of this feature matrix (over 350 million unique 31-mers), we build and use a sparse matrix representation. Our method, which we refer to as MTB++, leverages compact data structures and iterative methods to allow for the screening of all the 31-mers in the development of both LASSO LR and RF. MTB++ is able to achieve high discrimination (F-1 >80%) for the first-line antibiotics. Moreover, MTB++ had the highest F-1 score in all but three classes and was the most comprehensive since it had an F-1 score >75% in all but four (rare) antibiotic drugs. We use our feature selection to contextualize the 31-mers that are used for the prediction of phenotypic resistance, leading to some insights about sequence similarity to genes in MEGARes. Lastly, we give an estimate of the amount of data that is needed in order to provide accurate predictions.
引用
收藏
页码:i39 / i47
页数:9
相关论文
共 50 条
  • [41] *ACIDO-RESISTANCE DE MYCOBACTERIUM-TUBERCULOSIS ET HYDRAZIDE DE LACIDE ISONICOTINIQUE
    VIALLIER, J
    SERRE, H
    CAYRE, RM
    COMPTES RENDUS DES SEANCES DE LA SOCIETE DE BIOLOGIE ET DE SES FILIALES, 1953, 147 (15-1): : 1393 - 1395
  • [43] LA RESISTANCE ACQUISE DU MYCOBACTERIUM TUBERCULOSIS A LEGARD DE LISONICOTINHYDRAZIDE (INH)
    LEVADITI, C
    HENRYEVENO, J
    PRESSE MEDICALE, 1952, 60 (76): : 1633 - 1633
  • [44] Transcriptomic responses to antibiotic exposure in Mycobacterium tuberculosis
    Poonawala, Husain
    Zhang, Yu
    Kuchibhotla, Sravya
    Green, Anna G.
    Cirillo, Daniela Maria
    Di Marco, Federico
    Spitlaeri, Andrea
    Miotto, Paolo
    Farhat, Maha R.
    ANTIMICROBIAL AGENTS AND CHEMOTHERAPY, 2024, 68 (05)
  • [45] Quantitative measurement of antibiotic resistance in Mycobacterium tuberculosis reveals genetic determinants of resistance and susceptibility in a target gene approach
    Barilar, Ivan
    Battaglia, Simone
    Borroni, Emanuele
    Brandao, Angela Pires
    Brankin, Alice
    Cabibbe, Andrea Maurizio
    Carter, Joshua
    Chetty, Darren
    Cirillo, Daniela Maria
    Claxton, Pauline
    Clifton, David A.
    Cohen, Ted
    Coronel, Jorge
    Crook, Derrick W.
    Dreyer, Viola
    Earle, Sarah G.
    Escuyer, Vincent
    Ferrazoli, Lucilaine
    Fowler, Philip W.
    Gao, George Fu
    Gardy, Jennifer
    Gharbia, Saheer
    Ghisi, Kelen Teixeira
    Ghodousi, Arash
    Gibertoni Cruz, Ana Luiza
    Grandjean, Louis
    Grazian, Clara
    Groenheit, Ramona
    Guthrie, Jennifer L.
    He, Wencong
    Hoffmann, Harald
    Hoosdally, Sarah J.
    Hunt, Martin
    Iqbal, Zamin
    Ismail, Nazir Ahmed
    Jarrett, Lisa
    Joseph, Lavania
    Jou, Ruwen
    Kambli, Priti
    Khot, Rukhsar
    Knaggs, Jeff
    Koch, Anastasia
    Kohlerschmidt, Donna
    Kouchaki, Samaneh
    Lachapelle, Alexander S.
    Lalvani, Ajit
    Lapierre, Simon Grandjean
    Laurenson, Ian F.
    Letcher, Brice
    Lin, Wan-Hsuan
    NATURE COMMUNICATIONS, 2024, 15 (01)
  • [47] Reversing resistance to a tuberculosis antibiotic
    Everts, Sarah Y.
    CHEMICAL & ENGINEERING NEWS, 2017, 95 (12) : 5 - 5
  • [48] Structure-Based De Novo Design of Mycobacterium Tuberculosis VapC-Activating Stapled Peptides
    Kang, Sung-Min
    Moon, Heejo
    Han, Sang-Woo
    Kim, Hee
    Kim, Byeong Moon
    Lee, Bong-Jin
    ACS CHEMICAL BIOLOGY, 2020, 15 (09) : 2493 - 2498
  • [49] Structural investigations on orotate phosphoribosyltransferase from Mycobacterium tuberculosis, a key enzyme of the de novo pyrimidine biosynthesis
    Donini, Stefano
    Ferraris, Davide M.
    Miggiano, Riccardo
    Massarotti, Alberto
    Rizzi, Menico
    SCIENTIFIC REPORTS, 2017, 7
  • [50] De novo synthesized polyunsaturated fatty acids operate as both host immunomodulators and nutrients for Mycobacterium tuberculosis
    Laval, Thomas
    Pedro-Cos, Laura
    Malaga, Wladimir
    Guenin-Mace, Laure
    Pawlik, Alexandre
    Mayau, Veronique
    Yahia-Cherbal, Hanane
    Delos, Oceane
    Frigui, Wafa
    Bertrand-Michel, Justine
    Guilhot, Christophe
    Demangel, Caroline
    ELIFE, 2021, 10