ARTFEED — Contemporary Art Intelligence

Engineering Reliable Coding Agents: A Framework for Evaluation and Operation

ai-technology · 2026-08-17

A recent monograph available on arXiv (ID 2608.13867) posits that while AI coding agents are typically assessed as models, they are implemented as systems, with their reliability being contingent on the surrounding system components. This includes aspects such as harnessing, execution state, memory management, permissions, review interfaces, and resource allocation. The research integrates insights from 164 academic publications, 100 practitioner documents, 29 benchmark records, and 17 author-system case studies through a comprehensive multivocal review and targeted audits. The results reveal that many model failures stem from issues in the broader system, and enhancements at one level may not translate to overall performance improvements. The authors suggest a dependency chain approach for evaluation and operation, highlighting the significance of system-level factors over a model-focused perspective.

Key facts

  • The monograph is titled 'Engineering Reliable Coding Agents: Evaluating and Operating the System Around the Model'.
  • It is available on arXiv with ID 2608.13867.
  • It synthesizes 164 scholarly works, 100 practitioner records, 29 benchmark records, and 17 author-system case records.
  • The research uses a structured multivocal review, targeted update audits, software-engineering coverage analysis, and distributed-systems evidence synthesis.
  • The monograph argues that AI coding agents are evaluated as models but deployed as systems.
  • Reliability depends on the harness, execution state, retrieval, memory, state management, permissions, review interfaces, and resource allocation.
  • Many apparent model failures originate elsewhere in the system.
  • Improvements at one layer often fail to propagate to end-to-end outcomes.
  • Evaluation and operation are treated as a dependency chain.
  • Weaknesses in task construction, execution environments, and retrieval can undermine reliability.

Entities

Institutions

  • arXiv

Sources