Publication:

Contextual Multi-Armed Bandits Under Censorship

Loading...
Thumbnail Image

Files

AllenShen_SeniorThesis.pdf (1.68 MB)

Date

2026-04-09

Journal Title

Journal ISSN

Volume Title

Publisher

Research Projects

Organizational Units

Journal Issue

Access Restrictions

Abstract

In authoritarian regimes, centralized moderation algorithms and APIs enable censorship of political speech at industrial scales. Existing censorship evasion strategies---homophones, coded metaphors, manual rephrasing---are reactive, labor-intensive, and quickly neutralized as censors update their models. This thesis asks whether an automated agent can learn, through interaction with a live Chinese censorship system, to systematically transform censored text into forms that evade detection while preserving semantic meaning. We formalize this problem as a contextual multi-armed bandit, defining a 17-action space of meaning-preserving text transformations---spanning synonym substitution, phonetic rewriting, register shifting, sentence fragmentation, and character-level obfuscation---paired with a composite reward function that balances evasion success against semantic fidelity and edit cost. A frozen neural backbone, pre-trained on a censorship classification task, projects heterogeneous text features into a compact latent space over which a LinUCB agent learns context-dependent action-selection policies. Trained on 6,924 historically censored Sina Weibo posts from the Weiboscope corpus and evaluated against Alibaba Cloud's live Text Moderation API, the agent achieves a 67% (95% CI: [57.8%, 76.2%]) evasion rate on held-out data (an 8.4x improvement over a random baseline) and captures 74% of the oracle best-arm upper bound. Crucially, the agent exhibits emergent context-dependent routing: posts containing explicit sensitive keywords are directed toward character-level disruptions, while implicitly censored posts are routed to semantic restructuring strategies. To our knowledge, this work is among the first to cast censorship evasion as an online learning problem with a live oracle feedback loop, offering a principled, adaptive alternative to the artisanal evasion tactics that have long defined the censorship arms race.

Description

Type of resource

Princeton University Senior Theses

Keywords

Location

Citation