FMG-Bench: New Benchmark Evaluates AI's Theological Triage and Pastoral Guidance
The FMG-Bench (Faith & Moral Guidance Benchmark) has been launched to assess the responses of large language models (LLMs) to inquiries related to Christian faith, doctrine, and pastoral care. This benchmark, detailed in a paper on arXiv (ID: 2608.12324v1), seeks to fill a void in current evaluation techniques that overlook the complexity of such questions. It includes 120 scenarios in English, addressing fundamental Christian beliefs, differing viewpoints among faith traditions, prudential dilemmas requiring humility, and pastoral cases where human referral and safety are crucial. In its inaugural run, FMG-Bench v1 assessed 14 sophisticated models based on 8,792 scored responses, revealing that structured guidance enhances model performance by +3.96 points compared to unstructured behavior. The benchmark emphasizes the importance of AI recognizing when to defer to human insight in theological and pastoral matters.
Key facts
- FMG-Bench is a new benchmark for evaluating LLMs in Christian theological triage and pastoral guidance.
- It includes 120 scenarios in English.
- FMG-Bench v1 evaluated 14 advanced models across 8,792 scored responses.
- Three guided instruction settings were compared with raw model behavior.
- Structured harness improved performance by +3.96 points over raw model behavior.
- The benchmark covers core beliefs, disagreements, prudential issues, and pastoral safety.
- The paper is available on arXiv with ID 2608.12324v1.
Entities
Institutions
- arXiv