ARTFEED — Contemporary Art Intelligence

SQBench Benchmark Evaluates Language-Model Agents in Production Workflows

other · 2026-07-29

Researchers have unveiled SQBench, a new benchmark designed to assess how language-model agents perform in delivering production-oriented tasks. The initial version, SQBench v1.0, comprises 220 standardized tasks categorized into three levels: L1 for atomic capabilities, L2 for composite skills, and L3 for business scenarios. Each task necessitates that an agent handle input assets, utilize available tools, and generate a clearly defined output. The evaluation process first calculates functional Completion, followed by Risk Penalty and Performance, derived from independent triggers in a 10D Risk Matrix. A Strict Pass is defined as Completion = 1 and Risk Penalty = 0. The research assessed 27 model configurations using a unified protocol, with one execution for each configuration-task pair, achieving a maximum prespecified Weighted Pass@1 of 60.5%.

Key facts

  • SQBench is a benchmark for evaluating production-oriented task delivery by language-model agents.
  • SQBench v1.0 contains 220 standardized tasks.
  • Tasks are organized into L1 atomic capabilities, L2 composite skills, and L3 business scenarios.
  • Each task requires processing input assets, using tools, and producing a deliverable.
  • Evaluation computes Completion, then derives Risk Penalty and Performance from a 10D Risk Matrix.
  • A Strict Pass requires Completion = 1 and Risk Penalty = 0.
  • 27 model configurations were evaluated under a common protocol.
  • The highest prespecified Weighted Pass@1 was 60.5%.

Entities

Sources