Predictive Coding of Aligned Next-Generation Sequencing Data

被引:5
|
作者
Voges, Jan [1 ]
Munderloh, Marco [1 ]
Ostermann, Joern [1 ]
机构
[1] Leibniz Univ Hannover, TNT, Inst Informat Verarbeitung, Appelstr 9A, D-30167 Hannover, Germany
关键词
READ ALIGNMENT; COMPRESSION; FORMAT;
D O I
10.1109/DCC.2016.98
中图分类号
TP301 [理论、方法];
学科分类号
081202 ;
摘要
Due to novel high-throughput next-generation sequencing technologies, the sequencing of huge amounts of genetic information has become affordable. On account of this flood of data, IT costs have become a major obstacle compared to sequencing costs. High-performance compression of genomic data is required to reduce the storage size and transmission costs. The high coverage inherent in next-generation sequencing technologies produces highly redundant data. This paper describes a compression algorithm for aligned sequence reads. The proposed algorithm combines alignment information to implicitly assemble local parts of the donor genome in order to compress the sequence reads. In contrast to other algorithms, the proposed compressor does not need a reference to encode sequence reads. Compression is performed on-the-fly using solely a sliding window (i.e. a permanently updated short-time memory) as context for the prediction of sequence reads. The algorithm yields compression results on par or better than the state-of-the-art, compressing the data down to 1.9% of the original size at speeds of up to 60 MB/s and with a minute memory consumption of only several kilobytes, fitting in today's level 1 CPU caches.
引用
收藏
页码:241 / 250
页数:10
相关论文
共 50 条
  • [41] Computational classification of microRNAs in next-generation sequencing data
    Joshua Riback
    Artemis G. Hatzigeorgiou
    Martin Reczko
    [J]. Theoretical Chemistry Accounts, 2010, 125 : 637 - 642
  • [42] Model Testing of PluriTest with Next-Generation Sequencing Data
    Schulze, Markus
    Hoja, Sabine
    Winner, Beate
    Winkler, Juergen
    Edenhofer, Frank
    Riemenschneider, Markus J.
    [J]. STEM CELLS AND DEVELOPMENT, 2016, 25 (07) : 569 - 571
  • [43] NGSphy: phylogenomic simulation of next-generation sequencing data
    Escalona, Merly
    Rocha, Sara
    Posada, David
    [J]. BIOINFORMATICS, 2018, 34 (14) : 2506 - 2507
  • [44] Next-generation sequencing data analysis on cloud computing
    Taesoo Kwon
    Won Gi Yoo
    Won-Ja Lee
    Won Kim
    Dae-Won Kim
    [J]. Genes & Genomics, 2015, 37 : 489 - 501
  • [45] The Genome Assembly Model for Next-Generation Sequencing Data
    Wang, Yirong
    Wei, Chengdong
    Zhang, Xiaodong
    Cen, Tailin
    [J]. PROCEEDINGS OF THE 2017 INTERNATIONAL CONFERENCE ON APPLIED MATHEMATICS, MODELLING AND STATISTICS APPLICATION (AMMSA 2017), 2017, 141 : 97 - 101
  • [46] Next-Generation Anchor Based Phylogeny (NexABP): Constructing phylogeny from Next-generation sequencing data
    Tanmoy Roychowdhury
    Anchal Vishnoi
    Alok Bhattacharya
    [J]. Scientific Reports, 3
  • [47] Next-Generation Anchor Based Phylogeny (NexABP): Constructing phylogeny from Next-generation sequencing data
    Roychowdhury, Tanmoy
    Vishnoi, Anchal
    Bhattacharya, Alok
    [J]. SCIENTIFIC REPORTS, 2013, 3
  • [48] HUMAN DISEASE Next-generation sequencing of the next generation
    Burgess, Darren J.
    [J]. NATURE REVIEWS GENETICS, 2011, 12 (02) : 78 - 79
  • [49] Next-generation sequencing: The race is on
    von Bubnoff, Andreas
    [J]. CELL, 2008, 132 (05) : 721 - 723
  • [50] Combinatorics and next-generation sequencing
    Patterson, Nick
    Gabriel, Stacey
    [J]. NATURE BIOTECHNOLOGY, 2009, 27 (09) : 826 - 827