ARTFEED — Contemporary Art Intelligence

Scaffold Choice Can Shift Token Cost 40x in Coding Agents

ai-technology · 2026-07-29

A recent preprint on arXiv (2607.22585) indicates that the selection of scaffold—the software component that oversees tool usage, context, and termination criteria—can influence token usage by as much as 40 times for each task solved by coding agents, although pass-rate variations are limited to 0–8 percentage points. Researchers tested Qwen 3.6 Plus and MiniMax M2.5 using three open-source harnesses (Goose, OpenCode, OpenHands-SDK) on a carefully chosen subset of 50 tasks from Terminal-Bench Pro. Bootstrap confidence intervals (95%) within the same model showed no significant differences except for the largest disparity. Identified failure patterns—REASON for Goose, VERIFY/MAX_TURNS for OpenHands-SDK, and idle-loop/TIME for OpenCode—were consistent across models, suggesting biases at the harness level that are mostly independent of the model. The findings highlight that public leaderboards mix model and scaffold influences, advocating for standardized harness criteria to ensure equitable comparisons.

Key facts

  • arXiv preprint 2607.22585
  • Scaffold choice induces up to 40x difference in tokens per solved task
  • Pass-rate differences within 0-8 percentage points
  • Evaluated Qwen 3.6 Plus and MiniMax M2.5
  • Three open-source harnesses: Goose, OpenCode, OpenHands-SDK
  • Stratified 50-task subset of Terminal-Bench Pro
  • 95% paired-task bootstrap CIs include zero except for largest gap
  • Failure fingerprints replicate across models

Entities

Institutions

  • arXiv

Sources