Publication: On the Limitations of Concept Unlearning as a Mitigation Strategy for Child Image Generation
Files
Date
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Access Restrictions
Abstract
Text-to-image (T2I) models pose significant safety concerns due to the downstream tasks they enable. One such harmful downstream task is the production of non-consensual intimate imagery, which targets individuals of all ages. A particularly vulnerable group is children. In this context, perpetrators use different generative systems, driven by T2I models, to produce child sexual abuse material (CSAM). Prior works introduce gold-standard mitigation strategies like dataset filtering, but recent findings show these approaches are insufficient at preventing child generation. Consequently, attention has shifted to promising alternatives like concept unlearning that directly remove harmful concepts from pre-trained models.
Though recent works demonstrate unlearning methods' effectiveness on general object classes, artistic styles, and NSFW moderation, unlearning performance on safety-critical and semantically complex classes remains unexplored. We look at the extent to which existing unlearning methods can effectively remove concepts such as "child" from image generation. We evaluate multiple unlearning methods, including ESD, UCE, and negative prompting, across Stable Diffusion versions 1.4 and 2.1. With a diverse set of T2I prompts and a concept-detection-based model, we measure unlearning methods' effectiveness at mitigating child-related content generation pre- and post-unlearning.
We find that current unlearning methods are significantly less effective at removing safety-critical concepts than general object categories, and their performance degrades further with more detailed T2I prompts. These trends span across model variants, prompt types, and unlearned concepts. These factors contribute to our understanding of how existing unlearning methods perform on safety-critical concepts. Our extensive analysis highlights how existing unlearning approaches have fundamental limitations that hinder them from being a standalone solution for mitigating child generation.