OmegaUse-OfficeVal Benchmark Tests LLM Agents on Office Tasks
A new benchmark named OmegaUse-OfficeVal has been developed by researchers to assess large language model (LLM) agents on extended office-suite tasks grounded in economic principles. This benchmark comprises 100 tasks, all sourced from actual office-suite requests, which have been modified through a privacy-preserving method. Each task necessitates an average of 2.32 hours of human effort. Additionally, every task is associated with two economic indicators: the time taken by human labor and a proxy for task pricing, allowing for direct cost analysis between human work and LLM inference. The evaluation process is supported by code-based verifiers utilizing detailed rubrics. This research is documented in a paper available on arXiv (2607.27155).
Key facts
- OmegaUse-OfficeVal is a benchmark for LLM agents on office-suite tasks.
- It comprises 100 tasks from real office-suite requests.
- Tasks require an average of 2.32 hours of human labor.
- Each task has two economic signals: human labor time and task price proxy.
- Code-based verifiers from fine-grained rubrics are used for evaluation.
- The benchmark enables cost comparisons between humans and LLMs.
- The paper is available on arXiv with ID 2607.27155.
- The benchmark supports value-weighted evaluation.
Entities
Institutions
- arXiv