ARTFEED — Contemporary Art Intelligence

ESF-Bench Benchmark Reveals LLM Slot-Filling Limitations in Enterprise Settings

ai-technology · 2026-07-29

A new benchmark, ESF-Bench, has been released to evaluate large language models (LLMs) on slot-filling tasks in complex enterprise environments. The benchmark comprises 810 multi-turn samples and 6,530 slots across eight domains, curated from 57 challenging scenarios observed in real-world deployments. Testing revealed that GPT-OSS-120b, a state-of-the-art model, successfully extracted slots for only 20.7% of samples, highlighting significant limitations. The dataset, taxonomy, and evaluation code are publicly available on GitHub to support further research.

Key facts

  • ESF-Bench is a benchmark for enterprise slot filling.
  • It includes 810 multi-turn samples and 6,530 slots.
  • Covers 8 unique domains.
  • Based on a taxonomy of 57 challenging scenarios.
  • GPT-OSS-120b achieved only 20.7% success rate.
  • Benchmark dataset, taxonomy, and code are publicly released.
  • Focuses on real-world enterprise constraints and user behaviors.
  • Aims to improve LLM deployment in enterprise applications.

Entities

Institutions

  • arXiv
  • GitHub

Sources