Conrad: Gene prediction using conditional random fields

被引:37
|
作者
DeCaprio, David [1 ]
Vinson, Jade P.
Pearson, Matthew D.
Montgomery, Philip
Doherty, Matthew
Galagan, James E.
机构
[1] MIT, Broad Inst, Cambridge, MA 02142 USA
[2] Renaissance Technol LLC, New York, NY 11733 USA
关键词
D O I
10.1101/gr.6558107
中图分类号
Q5 [生物化学]; Q7 [分子生物学];
学科分类号
071010 ; 081704 ;
摘要
We present Conrad, the first comparative gene predictor based on semi-Markov conditional random fields (SMCRFs). Unlike the best standalone gene predictors, which are based on generalized hidden Markov models (GHMMs) and trained by maximum likelihood, Conrad is discriminatively trained to maximize annotation accuracy. In addition, unlike the best annotation pipelines, which rely on heuristic and ad hoc decision rules to combine standalone gene predictors with additional information such as ESTs and protein homology, Conrad encodes all sources of information as features and treats all features equally in the training and inference algorithms. Conrad outperforms the best standalone gene predictors in cross-validation and whole chromosome testing on two fungi with vastly different gene structures. The performance improvement arises from the SMCRF's discriminative training methods and their ability to easily incorporate diverse types of information by encoding them as feature functions. On Cryptococcus neoformans, configuring Conrad to reproduce the predictions of a two-species phylo-GHMM closely matches the performance of Twinscan. Enabling discriminative training increases performance, and adding new feature functions further increases performance, achieving a level of accuracy that is unprecedented for this organism. Similar results are obtained on Aspergillus nidulans comparing Conrad versus Fgenesh. SMCRFs are a promising framework for gene prediction because of their highly modular nature, simplifying the process of designing and testing potential indicators of gene structure. Conrad's implementation of SMCRFs advances the state of the art in gene prediction in fungi and provides a robust platform for both current application and future research.
引用
收藏
页码:1389 / 1398
页数:10
相关论文
共 50 条
  • [1] Personalized Driver Gene Prediction Using Graph Convolutional Networks with Conditional Random Fields
    Wei, Pi-Jing
    Zhu, An-Dong
    Cao, Ruifen
    Zheng, Chunhou
    [J]. BIOLOGY-BASEL, 2024, 13 (03):
  • [2] Punctuation Prediction for Vietnamese Texts Using Conditional Random Fields
    Pham, Quang H.
    Nguyen, Binh T.
    Nguyen Viet Cuong
    [J]. SOICT 2019: PROCEEDINGS OF THE TENTH INTERNATIONAL SYMPOSIUM ON INFORMATION AND COMMUNICATION TECHNOLOGY, 2019, : 322 - 327
  • [3] RECOGNITION OF GENE/PROTEIN NAMES USING CONDITIONAL RANDOM FIELDS
    Campos, David
    Matos, Sergio
    Oliveira, Jose Luis
    [J]. KDIR 2010: PROCEEDINGS OF THE INTERNATIONAL CONFERENCE ON KNOWLEDGE DISCOVERY AND INFORMATION RETRIEVAL, 2010, : 275 - 280
  • [4] Visual Webpage Block Importance Prediction Using Conditional Random Fields
    Tsai, Richard Tzong-Han
    Chiu, Borong
    Wu, Chi-En
    [J]. JOURNAL OF THE AMERICAN SOCIETY FOR INFORMATION SCIENCE AND TECHNOLOGY, 2011, 62 (11): : 2225 - 2235
  • [5] RNA secondary structure prediction using conditional random fields model
    Subpaiboonkit, Sitthichoke
    Thammarongtham, Chinae
    Cutler, Robert W.
    Chaijaruwanich, Jeerayut
    [J]. INTERNATIONAL JOURNAL OF DATA MINING AND BIOINFORMATICS, 2013, 7 (02) : 118 - 134
  • [6] Check-in Location Prediction using Wavelets and Conditional Random Fields
    Assam, Roland
    Seidl, Thomas
    [J]. 2014 IEEE INTERNATIONAL CONFERENCE ON DATA MINING (ICDM), 2014, : 713 - 718
  • [7] Letter-to-Sound Pronunciation Prediction Using Conditional Random Fields
    Wang, Dong
    King, Simon
    [J]. IEEE SIGNAL PROCESSING LETTERS, 2011, 18 (02) : 122 - 125
  • [8] Conditional random fields for transmembrane helix prediction
    Lukov, L
    Chawla, S
    Church, WB
    [J]. ADVANCES IN KNOWLEDGE DISCOVERY AND DATA MINING, PROCEEDINGS, 2005, 3518 : 155 - 161
  • [9] Identifying gene and protein mentions in text using conditional random fields
    McDonald, R
    Pereira, F
    [J]. BMC BIOINFORMATICS, 2005, 6 (Suppl 1)
  • [10] Identifying gene and protein mentions in text using conditional random fields
    Ryan McDonald
    Fernando Pereira
    [J]. BMC Bioinformatics, 6