Publication: GPUFlye: Balancing Efficiency and Accuracy in Long-Read Genome Assembly
Files
Date
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Access Restrictions
Abstract
Although de novo long-read genome assembly enables highly contiguous reconstruction of complex genomes, it remains computationally expensive due to the scale of modern sequencing datasets and the computational demands of key assembly stages. Flye is a widely used de novo long-read assembler that resolves repeats using a repeat graph representation followed by iterative polishing. Profiling across four genome datasets shows that Flye’s polishing stage accounts for an average of 37% of total assembly time, driven by repeated dynamic programming-based alignment and single-base edit evaluation across large numbers of bubbles and supporting branches. This work analyzes Flye’s polishing algorithm to identify performance bottlenecks and proposes GPUFlye, a CUDA-accelerated polisher that preserves Flye’s scoring and consensus mechanisms while exploiting fine-grained parallelism and high memory bandwidth on modern GPUs. Evaluated across four genome datasets, GPUFlye achieves polishing speedups between 1.22–5.95× and reduces total assembly time by up to 1.39×, with the NVIDIA L4 reducing full cloud assembly cost by up to 20% compared to a 16-thread CPU baseline.