Systematic Review Finds LLM Safety Alignment Weak in Low-Resource Languages
A recent literature review available on arXiv examines the challenges faced by large language models (LLMs) regarding safety alignment in multilingual and low-resource environments. Utilizing the PRISMA 2020 framework, researchers analyzed around 1,500 studies from platforms like Semantic Scholar and arXiv, ultimately narrowing it down to 50 pivotal works. The findings are categorized into four key themes: safety alignment techniques, risks in multilingual settings, evaluation standards, and cross-lingual adaptability. The authors underscore the inadequacy of translating English benchmarks, highlighting the urgent need for enhanced safety protocols tailored to diverse cultural contexts, particularly for languages with less support.
Key facts
- Systematic literature review on LLM safety alignment in low-resource languages
- Adopts PRISMA 2020 methodology
- Screened roughly 1,500 papers from Semantic Scholar, arXiv, and OpenAlex
- Selected 50 relevant studies for analysis
- Organized around four themes: safety alignment methods, multilingual safety risks, evaluation benchmarks, and cross-lingual transferability
- Proposes taxonomy of safety alignment approaches: data adaptation, objective optimization, mechanistic alignment
- Translated English benchmarks fail to sufficiently represent culturally rooted harms
- Published on arXiv with ID 2608.14626
Entities
Institutions
- Semantic Scholar
- arXiv
- OpenAlex