Publication: Adaptive Distributed Architecture for Ultra-Rapid Whole-Genome Sequencing
Files
Date
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Access Restrictions
Abstract
Motivated by the need to deliver rapid, actionable diagnoses in clinical settings where on-premise computing is limited, this work presents the design and development of a sequencing-aware distributed system for ultra-rapid genomic analysis, aimed at reducing cloud resource idleness and end-to-end latency in whole-genome sequencing (WGS) pipelines. Approximately 2 TB of raw nanopore signal data is produced within 90 minutes in sequencing workflows. However, current sequencing pipelines suffer from inherent stochastic variability, such as non-uniform flowcell performance, in which traditional static flowcell mapping designs result in inefficient cloud utilization and delays in diagnostic turnaround time (TAT).
This work analyzes the current ultra-rapid whole genome sequencing (urWGS) pipeline, identifying static flowcell-to-instance mapping as a key bottleneck that causes GPU idleness and necessitates manual intervention. To address this limitation, a flowcell-agnostic orchestration architecture is developed utilizing Ray, an open-source distributed computing framework that provides scheduling and autoscaling capabilities, to enable adaptive workload distribution and automated cloud resource management. The proposed system achieves efficient task allocation and elastic scaling across parallelized cloud resources, maximizing resource utilization while minimizing the number of active instances, effectively reducing both computation time and operational costs.