- Research Article
- 10.1093/gigascience/giag048
FEDRANN: effective long-read overlap detection based on dimensionality reduction and approximate nearest neighbors.
- May 08, 2026
- GigaScience
- Jia-Yuan Zhang + 14 more +14
Overlap detection is a key step in de novo genome assembly pipelines based on the Overlap-Layout-Consensus (OLC) paradigm. Existing methods for overlap detection either rely on heuristic seed-and-extension strategies or locality-sensitive hashing (LSH), both of which struggle to handle repetitive genomic regions and the computational burden of large-scale datasets. Here, we present FEDRANN, a novel strategy for overlap graph construction that integrates feature extraction, dimensionality reduction (DR), and approximate nearest neighbor (ANN) search. We find the pipeline combining inverse document frequency (IDF) transformation, sparse random projection (SRP), and NNDescent enables accurate detection of overlaps across diverse datasets. We developed an efficient open-source implementation of this pipeline named Fedrann (https://github.com/jzhang-dev/fedrann). Through systematic benchmarking on real long-read sequencing data, we demonstrate that Fedrann produces overlap graphs comparable to or better than those generated by existing state-of-the-art tools, including MECAT2, minimap2, and wtdbg2, while maintaining competitive runtime. By integrating Fedrann into the Shasta assembler, we successfully reconstructed human whole genomes, achieving high assembly contiguity and quality. Despite being implemented primarily in Python, Fedrann achieves performance parity with tools written in compiled languages by leveraging C-accelerated numerical libraries and optimized batch-based matrix operations. Our results suggest that the combination of dimensionality reduction and ANN techniques offers a robust, scalable framework for accurate overlap detection in long-read assembly and broader sequence similarity search tasks.
Read more