Publication:

Hardware-Enforced AI Kill-Switch via Trusted Monitor

Loading...
Thumbnail Image

Files

Forzani_Joey_Thesis.pdf (1.3 MB)

Date

2026-04-13

Journal Title

Journal ISSN

Volume Title

Publisher

Research Projects

Organizational Units

Journal Issue

Access Restrictions

Abstract

Software safeguards for AI systems operate within the same layer of abstraction as the models they constrain, leaving them vulnerable to circumvention by a sufficiently capable adversary. This thesis presents a hardware-enforced kill switch for GPU-accelerated AI systems that receives a trigger from a trusted external monitor and disables or throttles compute in response, without exposing any software interface through which the module can be bypassed. The trigger is a 64-byte ChaCha20-Poly1305 AEAD payload authenticated on-chip by a stored secret key and a replay counter, observed at a post-cache memory write bus tap that is invisible to the OS and executing model. I implement and verify this architecture in RTL for two GPU platforms, TinyGPU and VortexGPU, confirming correct authentication, replay rejection, and toggle semantics. I implement a steganographic payload delivery scheme that encodes payload bytes as natural-language sentences whose semantic embeddings carry the payload in their high-dimensional geometry, requiring full transformer inference to decode. Analysis of the steganographic scheme reveals that the carrier sentences are linguistically detectable by perplexity analysis, and that the decoding kernel's execution profile is distinguishable from normal inference, identifying the primary directions for future improvement. The cryptographic module contributes fewer than one part per million of GPU die area and approximately 1,376 cycles of authentication latency per trigger event, confirming that hardware-enforced kill-switch functionality is achievable at minimal silicon cost.

Description

Type of resource

Princeton University Senior Theses

Keywords

Location

Citation