HoosierHelp: New Benchmark Tests LLM Agents in Social Service Navigation
HoosierHelp has been developed by researchers as an interactive benchmark to assess large language model (LLM) agents in navigating social services. This benchmark is based on 3,971 publicly available social service resources from Indiana, USA. It replicates realistic user interactions by altering need structures, constraint satisfaction, and behaviors like impatience, rambling, unsupported requests, and self-contradiction. Agents must make structured resource-search requests, manage less-than-ideal interactions, and choose the final resources provided by the tool. Tests on 240 samples across seven LLMs indicate that current agents are significantly unreliable for social service navigation. This benchmark seeks to address the complexity of interactions and constraints that previous benchmarks have missed, offering a more accurate evaluation of LLM capabilities in this area. The research can be found on arXiv with the identifier 2608.09946.
Key facts
- HoosierHelp is an interactive benchmark for LLM agents in social service navigation.
- It is grounded in 3,971 Indiana public social service resources.
- Simulated users vary in need structure, constraint satisfiability, and behavior patterns.
- Behavior patterns include impatience, rambling, unsupported requests, and self-contradiction.
- Experiments were conducted on 240 samples across seven LLMs.
- Current LLM agents are substantially unreliable for social service navigation.
- The benchmark addresses interaction complexity and constraint-grounding demands.
- The study is available on arXiv (2608.09946).
Entities
Institutions
- arXiv
Locations
- Indiana
- United States