NGSpeciesID: DNA barcode and amplicon consensus generation from long-read sequencing data
Autor: | Marisa C. W. Lim, Stefan Prost, Kristoffer Sahlin |
---|---|
Rok vydání: | 2020 |
Předmět: |
0106 biological sciences
Computer science Computational biology 010603 evolutionary biology 01 natural sciences DNA sequencing 03 medical and health sciences lcsh:QH540-549.5 Preprocessor DNA barcoding sequence clustering Cluster analysis Ecology Evolution Behavior and Systematics 030304 developmental biology Nature and Landscape Conservation Sequence clustering Original Research 0303 health sciences Ecology amplicon sequencing business.industry Usability Amplicon third‐generation sequencing Filter (video) Scalability lcsh:Ecology Nanopore sequencing business |
Zdroj: | Ecology and Evolution Ecology and Evolution, Vol 11, Iss 3, Pp 1392-1398 (2021) |
Popis: | Third‐generation sequencing technologies, such as Oxford Nanopore Technologies (ONT) and Pacific Biosciences (PacBio), have gained popularity over the last years. These platforms can generate millions of long‐read sequences. This is not only advantageous for genome sequencing projects, but also advantageous for amplicon‐based high‐throughput sequencing experiments, such as DNA barcoding. However, the relatively high error rates associated with these technologies still pose challenges for generating high‐quality consensus sequences. Here, we present NGSpeciesID, a program which can generate highly accurate consensus sequences from long‐read amplicon sequencing technologies, including ONT and PacBio. The tool includes clustering of the reads to help filter out contaminants or reads with high error rates and employs polishing strategies specific to the appropriate sequencing platform. We show that NGSpeciesID produces consensus sequences with improved usability by minimizing preprocessing and software installation and scalability by enabling rapid processing of hundreds to thousands of samples, while maintaining similar consensus accuracy as current pipelines. Here we present NGSpeciesID, a program which can generate highly accurate consensus sequences from long‐read amplicon sequencing technologies, including ONT and PacBio. The tool includes clustering of the reads to help filter out contaminants or reads with high error rates and employs polishing strategies specific to the appropriate sequencing platform. |
Databáze: | OpenAIRE |
Externí odkaz: |