• Home
  • Search
  • A comparative study of k-spectrum-based error correction methods for next-generation sequencing data analysis
  • Cite Icon28
  • https://doi.org/10.1186/s40246-016-0068-0Copy DOI Icon

A comparative study of k-spectrum-based error correction methods for next-generation sequencing data analysis

Show More
  • Abstract
  • Highlights & Summary
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

BackgroundInnumerable opportunities for new genomic research have been stimulated by advancement in high-throughput next-generation sequencing (NGS). However, the pitfall of NGS data abundance is the complication of distinction between true biological variants and sequence error alterations during downstream analysis. Many error correction methods have been developed to correct erroneous NGS reads before further analysis, but independent evaluation of the impact of such dataset features as read length, genome size, and coverage depth on their performance is lacking. This comparative study aims to investigate the strength and weakness as well as limitations of some newest k-spectrum-based methods and to provide recommendations for users in selecting suitable methods with respect to specific NGS datasets.MethodsSix k-spectrum-based methods, i.e., Reptile, Musket, Bless, Bloocoo, Lighter, and Trowel, were compared using six simulated sets of paired-end Illumina sequencing data. These NGS datasets varied in coverage depth (10× to 120×), read length (36 to 100 bp), and genome size (4.6 to 143 MB). Error Correction Evaluation Toolkit (ECET) was employed to derive a suite of metrics (i.e., true positives, false positive, false negative, recall, precision, gain, and F-score) for assessing the correction quality of each method.ResultsResults from computational experiments indicate that Musket had the best overall performance across the spectra of examined variants reflected in the six datasets. The lowest accuracy of Musket (F-score = 0.81) occurred to a dataset with a medium read length (56 bp), a medium coverage (50×), and a small-sized genome (5.4 MB). The other five methods underperformed (F-score < 0.80) and/or failed to process one or more datasets.ConclusionsThis study demonstrates that various factors such as coverage depth, read length, and genome size may influence performance of individual k-spectrum-based error correction methods. Thus, efforts have to be paid in choosing appropriate methods for error correction of specific NGS datasets. Based on our comparative study, we recommend Musket as the top choice because of its consistently superior performance across all six testing datasets. Further extensive studies are warranted to assess these methods using experimental datasets generated by NGS platforms (e.g., 454, SOLiD, and Ion Torrent) under more diversified parameter settings (k-mer values and edit distances) and to compare them against other non-k-spectrum-based classes of error correction methods.

Similar Papers
  • Abstract
  • Citations1

Redefining “Gold Standard”: Ultra-Sensitive Characterization of Commercial DNA Standards with Duplex Sequencing

  • Nov 13, 2019
  • Blood
  • Jacob Higgins +4
  • Book Chapter
  • Citations3

Comprehensive Evaluation of Error-Correction Methodologies for Genome Sequencing Data

  • Mar 18, 2021
  • Bioinformatics
  • Yun Heo +3
  • Abstract
  • Citations1

Library Preparation Is the Major Factor Affecting Differences in Results of Immunoglobulin Gene Rearrangements Detection on Two Major Next-Generation Sequencing Platforms

  • Dec 03, 2015
  • Blood
  • Michaela Kotrová +21
  • Research Article

Abstract 5274: A novel, statistical-based method to determine RNA expression by next-generation sequencing in clinical FFPE samples

  • Jul 15, 2016
  • Cancer Research
  • Robert Shoemaker +8
  • PDF
  • Research Article
  • Citations2

English

  • Sep 26, 2018
  • African Journal of Biotechnology
  • Gatew H +1
  • Research Article
  • Citations2

Performance Comparison of NextSeq and Ion Proton Platforms for Molecular Diagnosis of Clinical Oncology

  • Jan 25, 2017
  • Tumori Journal
  • Fei Cao +7
  • Abstract

Leveraging The Old With The New: Exploring and Integrating Historic Microarray Studies With Next Generation Sequencing For Multiple Myeloma

  • Nov 15, 2013
  • Blood
  • Michael A Bauer +5
  • Research Article
  • Citations262

Statistical challenges associated with detecting copy number variations with next-generation sequencing

  • Aug 31, 2012
  • Bioinformatics
  • Shu Mei Teo +4
  • Conference Article
  • Citations5

Suffix-Tree Based Error Correction of NGS Reads Using Multiple Manifestations of an Error

  • Sep 22, 2013
  • Daniel M Savel +3
  • PDF
  • Research Article
  • Citations19

Pathosphere.org: pathogen detection and characterization through a web-based, open source informatics platform

  • Dec 01, 2015
  • BMC Bioinformatics
  • Andy Kilianski +23
  • Research Article
  • Citations30

Next generation sequencing and its application in deciphering head and neck cancer

  • Jan 15, 2014
  • Oral Oncology
  • Maryam Jessri +1
  • PDF
  • Discussion
  • Citations8

Dry Panels Supporting External Quality Assessment Programs for Next Generation Sequencing-Based HIV Drug Resistance Testing

  • Jun 20, 2020
  • Viruses
  • Marc Noguera-Julian +4
  • Abstract

Next Generation Amplicon Sequencing of Immunoglobulin Heavy Chain Gene Rearrangaments for Minimal Residual Disease (MRD) Stratification in Childhood Acute Lymphoblastic Leukemia (ALL): A Comparison with Classical qPCR-Based Technique

  • Dec 06, 2014
  • Blood
  • Michaela Kotrova +11
  • PDF
  • Research Article
  • Citations56

Targeted NGS Platforms for Genetic Screening and Gene Discovery in Primary Immunodeficiencies

  • Apr 11, 2019
  • Frontiers in Immunology
  • Cristina Cifaldi +39
  • Research Article
  • Citations52

Computational analysis of next generation sequencing data and its applications in clinical oncology

  • Jan 01, 2018
  • Informatics in Medicine Unlocked
  • Rucha M Wadapurkar +1
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.