Publication:

Dynamic Routing System for Accurate and Efficient Nanopore Basecalling

Loading...
Thumbnail Image

Files

Senior_Thesis_Final_Report_Sitao_Huang.pdf (11.54 MB)

Date

2026-04-13

Journal Title

Journal ISSN

Volume Title

Publisher

Research Projects

Organizational Units

Journal Issue

Access Restrictions

Abstract

Nanopore basecalling presents a tradeoff between accuracy and efficiency. Lightweight models provide high throughput but lower accuracy, whereas larger models achieve superior accuracy at substantially greater computational cost. This trade-off is particularly important for downstream genomic analyses, where localized basecalling errors can disproportionately affect tasks such as variant calling and metagenomic classification. This thesis investigates whether adaptive routing can bridge this gap by selectively assigning difficult signal regions to a stronger basecalling model while preserving the efficiency of a lighter model on easier regions.

The proposed framework introduces a lightweight pre-basecalling router that operates directly on raw nanopore signal. A key challenge is the absence of accessible ground-truth labels for third-party training. To address this, the thesis develops a self-alignment knowledge distillation framework in which the disagreement between a lightweight and a stronger basecaller is used as a pseudo-supervision signal. Router decisions are made at fine granularity using short overlapping signal windows, and are subsequently transformed into longer routed segments suitable for efficient basecalling without modifying Dorado's internal inference logic.

A compact convolutional--recurrent router is designed and trained as a binary classifier. The resulting router achieves a ROC--AUC of 0.877 on chunk-level routing. Experiments show that the router is highly lightweight relative to the basecalling models, both in parameter footprint and in normalized compute cost, so its overhead is negligible under the proposed cost model. On whole-read basecalling, the routed system exhibits the expected accuracy--efficiency trade-off, with the strongest gains occurring when the router selectively escalates the most difficult regions first. On downstream evaluation, the routed pipeline substantially narrows the gap between HAC and SUP in variant calling, especially for indel recall and precision, while also improving metagenomic abundance estimation and reducing the proportion of unclassified reads.

Overall, this thesis demonstrates that adaptive routing is a viable direction for nanopore basecalling. It shows that substantial downstream benefits can be obtained without globally paying the cost of the strongest basecalling model, and establishes a practical framework for learning routing decisions from unlabeled nanopore data.

Description

Type of resource

Princeton University Senior Theses

Keywords

Location

Citation