ESF-Bench Benchmark Reveals LLM Slot-Filling Limitations in Enterprise Settings
A new benchmark, ESF-Bench, has been released to evaluate large language models (LLMs) on slot-filling tasks in complex enterprise environments. The benchmark comprises 810 multi-turn samples and 6,530 slots across eight domains, curated from 57 challenging scenarios observed in real-world deployments. Testing revealed that GPT-OSS-120b, a state-of-the-art model, successfully extracted slots for only 20.7% of samples, highlighting significant limitations. The dataset, taxonomy, and evaluation code are publicly available on GitHub to support further research.
Key facts
- ESF-Bench is a benchmark for enterprise slot filling.
- It includes 810 multi-turn samples and 6,530 slots.
- Covers 8 unique domains.
- Based on a taxonomy of 57 challenging scenarios.
- GPT-OSS-120b achieved only 20.7% success rate.
- Benchmark dataset, taxonomy, and code are publicly released.
- Focuses on real-world enterprise constraints and user behaviors.
- Aims to improve LLM deployment in enterprise applications.
Entities
Institutions
- arXiv
- GitHub