RNA-QC-chain: comprehensive and fast quality control for RNA-Seq data

被引:38
|
作者
Zhou, Qian [1 ,2 ]
Su, Xiaoquan [3 ,4 ,5 ]
Jing, Gongchao [3 ,4 ]
Chen, Songlin [1 ,2 ]
Ning, Kang [6 ]
机构
[1] Chinese Acad Fishery Sci, Key Lab Sustainable Dev Marine Fisheries, Minist Agr, Yellow Sea Fisheries Res Inst, Qingdao 266071, Shandong, Peoples R China
[2] Qingdao Natl Lab Marine Sci & Technol, Lab Marine Fisheries Sci & Food Prod Proc, Qingdao 266071, Shandong, Peoples R China
[3] Chinese Acad Sci, Qingdao Inst Bioenergy & Bioproc Technol, CAS Key Lab Biofuels, Shandong Key Lab Energy Genet, Qingdao 266101, Shandong, Peoples R China
[4] Chinese Acad Sci, Qingdao Inst Bioenergy & Bioproc Technol, Single Cell Ctr, Qingdao 266101, Shandong, Peoples R China
[5] Univ Chinese Acad Sci, Beijing 100049, Peoples R China
[6] Huazhong Univ Sci & Technol, Coll Life Sci & Technol, Hubei Key Lab Bioinformat & Mol Imaging, Minist Educ,Key Lab Mol Biophys,Dept Bioinformat, Wuhan 430074, Hubei, Peoples R China
来源
BMC GENOMICS | 2018年 / 19卷
基金
中国国家自然科学基金;
关键词
Quality control; RNA-Seq; Contamination identification; Alignment statistics; Parallel computing;
D O I
10.1186/s12864-018-4503-6
中图分类号
Q81 [生物工程学(生物技术)]; Q93 [微生物学];
学科分类号
071005 ; 0836 ; 090102 ; 100705 ;
摘要
Background: RNA-Seq has become one of the most widely used applications based on next-generation sequencing technology. However, raw RNA-Seq data may have quality issues, which can significantly distort analytical results and lead to erroneous conclusions. Therefore, the raw data must be subjected to vigorous quality control (QC) procedures before downstream analysis. Currently, an accurate and complete QC of RNA-Seq data requires of a suite of different QC tools used consecutively, which is inefficient in terms of usability, running time, file usage, and interpretability of the results. Results: We developed a comprehensive, fast and easy-to-use QC pipeline for RNA-Seq data, RNA-QC-Chain, which involves three steps: (1) sequencing-quality assessment and trimming; (2) internal (ribosomal RNAs) and external (reads from foreign species) contamination filtering; (3) alignment statistics reporting (such as read number, alignment coverage, sequencing depth and pair-end read mapping information). This package was developed based on our previously reported tool for general QC of next-generation sequencing (NGS) data called QC-Chain, with extensions specifically designed for RNA-Seq data. It has several features that are not available yet in other QC tools for RNA-Seq data, such as RNA sequence trimming, automatic rRNA detection and automatic contaminating species identification. The three QC steps can run either sequentially or independently, enabling RNA-QC-Chain as a comprehensive package with high flexibility and usability. Moreover, parallel computing and optimizations are embedded in most of the QC procedures, providing a superior efficiency. The performance of RNA-QC-Chain has been evaluated with different types of datasets, including an in-house sequencing data, a semi-simulated data, and two real datasets downloaded from public database. Comparisons of RNA-QC-Chain with other QC tools have manifested its superiorities in both function versatility and processing speed. Conclusions: We present here a tool, RNA-QC-Chain, which can be used to comprehensively resolve the quality control processes of RNA-Seq data effectively and efficiently.
引用
收藏
页数:10
相关论文
共 50 条
  • [1] RNA-QC-chain: comprehensive and fast quality control for RNA-Seq data
    Qian Zhou
    Xiaoquan Su
    Gongchao Jing
    Songlin Chen
    Kang Ning
    BMC Genomics, 19
  • [2] QoRTs: a comprehensive toolset for quality control and data processing of RNA-Seq experiments
    Stephen W. Hartley
    James C. Mullikin
    BMC Bioinformatics, 16
  • [3] QoRTs: a comprehensive toolset for quality control and data processing of RNA-Seq experiments
    Hartley, Stephen W.
    Mullikin, James C.
    BMC BIOINFORMATICS, 2015, 16
  • [4] A comprehensive review on RNA-seq data analysis
    Zhang, Li
    Liu, Xuejun
    Transactions of Nanjing University of Aeronautics and Astronautics, 2016, 33 (03) : 339 - 361
  • [5] A Comprehensive Review on RNA-seq Data Analysis
    Zhang Li
    Liu Xuejun
    TransactionsofNanjingUniversityofAeronauticsandAstronautics, 2016, 33 (03) : 339 - 361
  • [6] The impact of quality control in RNA-seq experiments
    Merino, Gabriela A.
    Fresno, Cristobal
    Netto, Frederico
    Netto, Emmanuel Dias
    Pratto, Laura
    Fernandez, Elmer A.
    20TH ARGENTINEAN BIOENGINEERING SOCIETY CONGRESS (XX ARGENTINE BIOENGINEERING CONGRESS AND IX CONFERENCE OF CLINICAL ENGINEERING), (SABI 2015), 2016, 705
  • [8] RSeQC: quality control of RNA-seq experiments
    Wang, Liguo
    Wang, Shengqin
    Li, Wei
    BIOINFORMATICS, 2012, 28 (16) : 2184 - 2185
  • [9] A comprehensive simulation study on classification of RNA-Seq data
    Zararsiz, Gokmen
    Goksuluk, Dincer
    Korkmaz, Selcuk
    Eldem, Vahap
    Zararsiz, Gozde Erturk
    Duru, Izzet Parug
    Ozturk, Ahmet
    PLOS ONE, 2017, 12 (08):
  • [10] A comprehensive workflow for optimizing RNA-seq data analysis
    Jiang, Gao
    Zheng, Juan-Yu
    Ren, Shu-Ning
    Yin, Weilun
    Xia, Xinli
    Li, Yun
    Wang, Hou-Ling
    BMC GENOMICS, 2024, 25 (01):