Test-Time Safety Mechanism for Text-to-Image Diffusion Models
A recent research paper introduces a test-time strategy aimed at enhancing safety in text-to-image diffusion models, specifically tackling the issue of creating banned content like nudity and protected intellectual property from both benign and adversarial prompts. In contrast to unlearning methods based on training, which can be costly and may lead to significant loss of general capabilities, this method utilizes intermediate clean image estimates during generation. It incorporates a sparse margin objective to identify prohibited concepts. Upon detecting a violation, the system promptly optimizes a structured low-rank residual in the text-conditioning space using truncated backpropagation. This design preserves weight detection, ensuring that non-violating inferences remain unaffected. The paper can be found on arXiv with the identifier 2608.03284.
Key facts
- Paper arXiv:2608.03284 proposes test-time safety mechanism for text-to-image diffusion models.
- Method uses intermediate clean image estimates for detection of prohibited content.
- Sparse margin objective is employed to detect prohibited concepts.
- Intervention optimizes a structured low-rank residual in text-conditioning space via truncated backpropagation.
- Approach is weight-preserving and does not affect non-violating inference.
- Addresses limitations of training-based unlearning methods.
- Targets prohibited content like nudity and protected intellectual property.
- Paper is a cross-announcement on arXiv.
Entities
Institutions
- arXiv