ARTFEED — Contemporary Art Intelligence

OmegaUse-OfficeVal Benchmark Tests LLM Agents on Office Tasks

ai-technology · 2026-07-30

A new benchmark named OmegaUse-OfficeVal has been developed by researchers to assess large language model (LLM) agents on extended office-suite tasks grounded in economic principles. This benchmark comprises 100 tasks, all sourced from actual office-suite requests, which have been modified through a privacy-preserving method. Each task necessitates an average of 2.32 hours of human effort. Additionally, every task is associated with two economic indicators: the time taken by human labor and a proxy for task pricing, allowing for direct cost analysis between human work and LLM inference. The evaluation process is supported by code-based verifiers utilizing detailed rubrics. This research is documented in a paper available on arXiv (2607.27155).

Key facts

  • OmegaUse-OfficeVal is a benchmark for LLM agents on office-suite tasks.
  • It comprises 100 tasks from real office-suite requests.
  • Tasks require an average of 2.32 hours of human labor.
  • Each task has two economic signals: human labor time and task price proxy.
  • Code-based verifiers from fine-grained rubrics are used for evaluation.
  • The benchmark enables cost comparisons between humans and LLMs.
  • The paper is available on arXiv with ID 2607.27155.
  • The benchmark supports value-weighted evaluation.

Entities

Institutions

  • arXiv

Sources