FairFund-Bench: New Benchmark to Evaluate Bias in LLM Resource Allocation
The newly launched benchmark, FairFund-Bench, aims to systematically assess distributive bias in large language models (LLMs) regarding the allocation of limited resources. This benchmark tackles the inconsistencies found in earlier audits of LLMs, which have shown both favorable and unfavorable treatment of women and ethnic minorities, even within the same models. FairFund-Bench modifies essential aspects of audit designs, including the evaluation task (rating, ranking, or allocation), the comparison context (single or multi-stimulus), and the level of transparency (transparent or disguised). It consists of 600 financial assistance requests derived from human-written templates, aligned with 1.3 million actual GoFundMe campaigns, spanning three domains, four racial categories, two gender categories, and five causal need framings. The goal is to enhance the reliability of bias assessments in LLM-based resource distribution, particularly in financial aid contexts.
Key facts
- FairFund-Bench is a new benchmark for evaluating distributive bias in LLMs.
- It addresses inconsistencies in previous LLM audits regarding bias towards women and ethnic minorities.
- The benchmark varies evaluation task, comparison context, and transparency of audits.
- It includes 600 requests for financial assistance based on human-authored templates.
- The templates were calibrated against 1.3 million real GoFundMe campaigns.
- The benchmark covers three domains, four race categories, two gender categories, and five causal framings of need.
- The research is available on arXiv with identifier 2607.28934.
- The goal is to improve assessment of bias in LLM resource allocation.
Entities
Institutions
- arXiv