Accurate gene consensus at low nanopore coverage

Abstract Background Nanopore technologies allow high-throughput sequencing of long strands of DNA at the cost of a relatively large error rate. This limits its use in the reading of amplicon libraries in which there are only a few mutations per variant and therefore they are easily confused with the...

Ausführliche Beschreibung

Gespeichert in:
Bibliographische Detailangaben
Veröffentlicht in:Gigascience 2022-11, Vol.11
Hauptverfasser: Espada, Rocío, Zarevski, Nikola, Dramé-Maigné, Adèle, Rondelez, Yannick
Format: Artikel
Sprache:eng
Schlagworte:
Online-Zugang:Volltext
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
Beschreibung
Zusammenfassung:Abstract Background Nanopore technologies allow high-throughput sequencing of long strands of DNA at the cost of a relatively large error rate. This limits its use in the reading of amplicon libraries in which there are only a few mutations per variant and therefore they are easily confused with the sequencing noise. Consensus calling strategies reduce the error but sacrifice part of the throughput on reading typically 30 to 100 times each member of the library. Findings In this work, we introduce SINGLe (SNPs In Nanopore reads of Gene Libraries), an error correction method to reduce the noise in nanopore reads of amplicons containing point variations. SINGLe exploits that in an amplicon library, all reads are very similar to a wild-type sequence from which it is possible to experimentally characterize the position-specific systematic sequencing error pattern. Then, it uses this information to reweight the confidence given to nucleotides that do not match the wild-type in individual variant reads and incorporates it on the consensus calculation. Conclusions We tested SINGLe in a mutagenic library of the KlenTaq polymerase gene, where the true mutation rate was below the sequencing noise. We observed that contrary to other methods, SINGLe compensates for the systematic errors made by the basecallers. Consequently, SINGLe converges to the true sequence using as little as 5 reads per variant, fewer than the other available methods.
ISSN:2047-217X
2047-217X
DOI:10.1093/gigascience/giac102