Publication: Chip of Theseus: Redundant SoC Architecture for Extending System Longevity
Files
Date
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Access Restrictions
Abstract
The combination of slowing of transistor scaling and diminishing performance gains from new hardware generations, the acceleration of device aging, the continued rise in electronic waste, and the fact that everyday workloads no longer require frequent hardware upgrades, motivates chip designs that prioritize longevity and reduce the need for frequent device replacement. This work explores microarchitectural strategies for enabling long-lived processors through fine-grained hardware redundancy and graceful performance degradation.
We explore the feasibility of structural duplication and propose and implement three distinct redundancy architectures targeting execution units within multi-core in-order and out-of-order RISC-V cores using the Chipyard framework. These include local per-core multiplexing with cold-spare units, a shared redundancy accelerator amortized across multiple cores, and a TileLink-connected redundancy bank accessed via memory-mapped operations. The strategies represent differing tradeoffs across latency, area overhead, and system complexity. To evaluate these approaches, we develop configurable SoC designs, perform simulation and synthesis, and analyze performance retention under degraded conditions and different wear-out models.
In addition, we introduce a redundancy decision-making model grounded in realistic wear-out distributions, enabling informed architectural decisions based on expected failure rates, area constraints, and performance requirements. Our results show that while local redundancy minimizes latency overhead, shared approaches significantly reduce area cost at the expense of performance, particularly for latency-sensitive operations. Overall, this work demonstrates the feasibility of incorporating fine-grained redundancy into modern processors and provides a framework and starting place for designing resilient SoCs capable of maintaining correct operation under permanent faults, extending device lifetime.