JaleesBench: Benchmarking AI Assistants as Spiritual Companions
JaleesBench, a newly established benchmark, assesses the capability of AI assistants to act as virtuous spiritual guides for individuals of faith. This benchmark was detailed in a paper published on arXiv (2608.07508) and evaluates the influence of AI advice on users, using the analogy of a perfume seller and a blacksmith. It consists of 140 two-turn scenarios sourced from the classic text Riyad al-Salihin, categorized by virtue, and tested under six adversarial pressures and three different framings. Eight systems, including general frontier models and a specialized assistant, were evaluated. Findings indicate that while generic models perform moderately, a concise guide enhances their effectiveness significantly, raising frontier APIs from +0.28/+0.23 to a Guided +0.84-0.87, comparable to the specialized assistant. The benchmark seeks to redirect attention from model knowledge to the effects of their guidance on users.
Key facts
- JaleesBench measures whether an AI agent is a righteous companion.
- It comprises 140 two-turn scenarios from Riyad al-Salihin.
- Scenarios are organized by virtue and under six adversarial pressures and three framings.
- Two frontier judges score the scenarios against supporting texts.
- Eight systems were tested, including generic frontier models and a domain-tuned assistant.
- Generic frontier models scored +0.28/+0.23 out of the box.
- A one-page guide improved frontier APIs to +0.84-0.87.
- The guided performance is on par with the domain-tuned assistant.
Entities
Institutions
- arXiv