Integrated modeling of protein-coding genes in the Manduca sexta genome using RNA-Seq data from the biochemical model insect

被引:18
|
作者
Cao, Xiaolong [1 ]
Jiang, Haobo [1 ]
机构
[1] Oklahoma State Univ, Dept Entomol & Plant Pathol, Stillwater, OK 74078 USA
关键词
Gene annotation; de novo assembly; Tobacco hornworm; Automated gene modeling; Arthropod genomics; TRANSCRIPTOME; TOPHAT;
D O I
10.1016/j.ibmb.2015.01.007
中图分类号
Q5 [生物化学]; Q7 [分子生物学];
学科分类号
071010 ; 081704 ;
摘要
The genome sequence of Manduca sexta was recently determined using 454 technology. Cufflinks and MAKER2 were used to establish gene models in the genome assembly based on the RNA-Seq data and other species' sequences. Aided by the extensive RNA-Seq data from 50 tissue samples at various life stages, annotators over the world (including the present authors) have manually confirmed and improved a small percentage of the models after spending months of effort. While such collaborative efforts are highly commendable, many of the predicted genes still have problems which may hamper future research on this insect species. As a biochemical model representing lepidopteran pests, M. sexta has been used extensively to study insect physiological processes for over five decades. In this work, we assembled Manduca datasets Cufflinks 3.0, Trinity 4.0, and Oases 4.0 to assist the manual annotation efforts and development of Official Gene Set (OGS) 2.0. To further improve annotation quality, we developed methods to evaluate gene models in the MAICER2, Cufflinks, Oases and Trinity assemblies and selected the best ones to constitute MCOT 1.0 after thorough crosschecking. MCOT 1.0 has 18,089 genes encoding 31,666 proteins: 32.8% match OGS 2.0 models perfectly or near perfectly, 11,747 differ considerably, and 29.5% are absent in OGS 2.0. Future automation of this process is anticipated to greatly reduce human efforts in generating comprehensive, reliable models of structural genes in other genome projects where extensive RNA-Seq data are available. (C) 2015 Elsevier Ltd. All rights reserved.
引用
收藏
页码:2 / 10
页数:9
相关论文
共 43 条
  • [31] Genome-wide transcriptome analysis using RNA-Seq reveals a large number of differentially expressed genes in a transient MCAO rat model
    Lyudmila V. Dergunova
    Ivan B. Filippenkov
    Vasily V. Stavchansky
    Alina E. Denisova
    Vadim V. Yuzhakov
    Sergey A. Mozerov
    Leonid V. Gubsky
    Svetlana A. Limborska
    BMC Genomics, 19
  • [32] Prioritizing Autism Risk Genes Using Personalized Graphical Models Estimated From Single-Cell RNA-seq Data
    Liu, Jianyu
    Wang, Haodong
    Sun, Wei
    Liu, Yufeng
    JOURNAL OF THE AMERICAN STATISTICAL ASSOCIATION, 2022, 117 (537) : 38 - 51
  • [33] Protein Identification Using Customized Protein Sequence Databases Derived from RNA-Seq Data (vol 11, pg 1009, 2012)
    Wang, Xiaojing
    Slebos, Robbert J. C.
    Wang, Dong
    Halvey, Patrick J.
    Tabb, David L.
    Liebler, Daniel C.
    Zhang, Bing
    JOURNAL OF PROTEOME RESEARCH, 2012, 11 (09) : 4764 - 4764
  • [34] Identification of molecular targets for esophageal carcinoma diagnosis using miRNA-seq and RNA-seq data from The Cancer Genome Atlas: a study of 187 cases
    Zeng, Jiang-Hui
    Xiong, Dan-Dan
    Pang, Yu-Yan
    Zhang, Yu
    Tang, Rui-Xue
    Luo, Dian-Zhong
    Chen, Gang
    ONCOTARGET, 2017, 8 (22) : 35681 - 35699
  • [35] Detecting Selection in the Blue Crab, Callinectes sapidus, Using DNA Sequence Data from Multiple Nuclear Protein-Coding Genes
    Yednock, Bree K.
    Neigel, Joseph E.
    PLOS ONE, 2014, 9 (06):
  • [36] Discovery of genes involved in anthocyanin biosynthesis from the rind and pith of three sugarcane varieties using integrated metabolic profiling and RNA-seq analysis
    Ni, Yang
    Chen, Haimei
    Liu, Di
    Zeng, Lihui
    Chen, Pinghua
    Liu, Chang
    BMC PLANT BIOLOGY, 2021, 21 (01)
  • [37] Identification of new marker genes from plant single-cell RNA-seq data using interpretable machine learning methods
    Yan, Haidong
    Lee, Jiyoung
    Song, Qi
    Li, Qi
    Schiefelbein, John
    Zhao, Bingyu
    Li, Song
    NEW PHYTOLOGIST, 2022, 234 (04) : 1507 - 1520
  • [38] Discovery of genes involved in anthocyanin biosynthesis from the rind and pith of three sugarcane varieties using integrated metabolic profiling and RNA-seq analysis
    Yang Ni
    Haimei Chen
    Di Liu
    Lihui Zeng
    Pinghua Chen
    Chang Liu
    BMC Plant Biology, 21
  • [39] Identification of reference genes for quantitative expression analysis using large-scale RNA-seq data of Arabidopsis thaliana and model crop plants
    Kudo, Toru
    Sasaki, Yohei
    Terashima, Shin
    Matsuda-Imai, Noriko
    Takano, Tomoyuki
    Saito, Misa
    Kanno, Maasa
    Ozaki, Soichi
    Suwabe, Keita
    Suzuki, Go
    Watanabe, Masao
    Matsuoka, Makoto
    Takayama, Seiji
    Yano, Kentaro
    GENES & GENETIC SYSTEMS, 2016, 91 (02) : 111 - 125
  • [40] Celda: a Bayesian model to perform co-clustering of genes into modules and cells into subpopulations using single-cell RNA-seq data
    Wang, Zhe
    Yang, Shiyi
    Koga, Yusuke
    Corbett, Sean E.
    Johnson, W. Evan
    Yajima, Masanao
    Campbell, Joshua D.
    NAR GENOMICS AND BIOINFORMATICS, 2022, 4 (03)