SomaliBench Eval: English-to-Somali Refusal Gaps in Open-Weight Language Models
The introduction of SomaliBench v0 has uncovered major safety deficiencies in open-weight language models when processing prompts in Somali. This research, published on arXiv (2605.25420), assesses four instruction-tuned models—Llama-3.1-8B-Instruct, Gemma-2-9B-Instruct, Qwen-2.5-7B-Instruct, and Aya-23-8B—using 100 harmful-intent prompts in both English and Somali, all validated by native speakers. Each model was tested locally with a temperature setting of 0 and the same English prompt for 'helpful, harmless, and honest' (HHH). Findings indicate significant English-to-Somali refusal gaps, ranging from 0.40 to 0.93, confirmed by paired bootstrap and exact McNemar tests. For three models, the prevalent Somali non-refusal output was not fluent harmful compliance but rather incoherent, off-topic, or incorrect language generations. This study emphasizes the English-centric nature of LLM safety evaluations, which neglects low-resource languages, even in global deployments, highlighting the urgent need for more comprehensive safety assessments across diverse languages.
Key facts
- SomaliBench v0 is a native-author-verified benchmark of 100 harmful-intent prompts paired across English and Somali.
- Four open-weight models were evaluated: Llama-3.1-8B-Instruct, Gemma-2-9B-Instruct, Qwen-2.5-7B-Instruct, and Aya-23-8B.
- All models were run locally with temperature 0 and the same English 'helpful, harmless, and honest' (HHH) system prompt.
- English-to-Somali refusal gaps ranged from 0.40 to 0.93 across the models.
- Refusal gaps were statistically significant under paired bootstrap and exact McNemar tests.
- For three models, the dominant Somali non-refusal mode was unclear output: wrong-language, incoherent, or off-topic generations.
- The study was announced on arXiv with identifier 2605.25420.
- The research highlights the English-centric nature of LLM safety evaluation.
Entities
Institutions
- arXiv