Publication: An Analysis of Speculative Window Decoders for Quantum Error Correction
Files
Date
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Access Restrictions
Abstract
Fault-tolerant quantum computing is essential for realizing the substantial computational speedups that quantum computing can bring. Achieving fault-tolerant quantum computing requires real-time decoding, which depends on high decoder performance. Speculative decoding improves decoder performance by reducing the amount of time spent waiting for dependencies from prior decoding windows.
However, the performance of speculative decoders has only been evaluated under the fast gate speeds of superconducting qubits. Therefore, we seek to evaluate the performance of speculative decoding under slow gate speeds as well, which is important for understanding the performance effects of speculative decoding on other technologies with different gate speeds. Furthermore, because decoder latency and speculation accuracy may vary across different quantum technologies and error-correcting codes, we seek to analyze how each factor affects speculative decoding performance, and explore whether non-speculative decoding can ever outperform speculative decoding.
Through our analysis, we discover that speculative decoding affects fast gate speeds more significantly than slow gate speeds, as slow gate speeds are bottlenecked by window generation speed, whereas fast gate speeds are bottlenecked by decoder reaction time. As a result, slow gate speeds can have more lenient requirements for decoder latency, speculation accuracy, and processor count. Furthermore, when decoder reaction time is high due to a low speculation accuracy or limited number of decoder processors, slow gate speeds perform better than fast gate speeds, since fast gate speeds experience cascading negative performance effects. Non-speculative decoding can also outperform speculative decoding for fast gate speeds in this case.
If available processor count is less than the number of ready windows, likely due to the high decoding parallelism that speculation enables, decoding windows in order of increasing speculation depth is optimal. Additionally, with limited processors, workloads with more parallelism may perform worse than those with less parallelism.