Community-Specific Toxicity Detection Needed for Text-to-Image AI
A new paper on arXiv (2607.24898) argues that current toxicity detectors for text-to-image generation fail marginalized communities. These detectors adopt a one-size-fits-all approach, but approximately 35% of images labeled safe are considered harmful by disability communities. The authors propose community-specific toxicity detection (CTD) and collaborate with disability experts to develop safety guidelines for dwarfism and blind/low vision communities. Using a dataset of 2,400 annotated T2I-generated images, they show that both large vision-language models and general-purpose detectors catastrophically fail in zero-shot settings, with F1 scores of 0.32 and 0.37—lower than random guessing. Prompt-based adaptation shows promise.
Key facts
- arXiv:2607.24898v1
- State-of-the-art toxicity detectors use a universal model
- 35% of images labeled safe are harmful to disability communities
- Paper argues for community-specific toxicity detection (CTD)
- Collaboration with disability experts for dwarfism and blind/low vision
- Dataset of 2,400 annotated T2I-generated images
- F1 scores 0.32 and 0.37 for zero-shot detection
- Prompt-based adaptation shows promise
Entities
Institutions
- arXiv