ARTFEED — Contemporary Art Intelligence

Active-SWE: Benchmarking Coding Agents for Proactive Bug Fixing

ai-technology · 2026-08-06

Active-SWE is a new benchmark introduced in a research paper on arXiv (2608.04682) that evaluates coding agents powered by large language models (LLMs) on their ability to proactively discover and fix multiple bugs without relying on issue reports. The benchmark covers 1,663 tasks across six bug categories and eight programming languages, shifting the focus from reactive bug fixing (where detailed issue reports are provided) to proactive bug fixing. The construction of Active-SWE involves a novel difficulty-aware task selection method, as described in the abstract. This work addresses a practical gap in software engineering, where high-quality issue reports are not always available. The benchmark aims to enable more in-depth evaluation of coding agents by expanding the scope to multiple-bug fixing and potential bug discovery scenarios. The paper is authored by researchers and was announced as a cross-type submission on arXiv.

Key facts

  • Active-SWE is a benchmark for evaluating coding agents on proactive bug fixing.
  • It covers 1,663 tasks across six bug categories and eight languages.
  • The benchmark shifts from reactive to proactive bug fixing.
  • It expands evaluation to multiple-bug fixing and potential bug discovery.
  • The construction uses a difficulty-aware task selection method.
  • The paper is available on arXiv with ID 2608.04682.
  • The announcement type is 'cross'.
  • The benchmark addresses the lack of issue reports in real-world scenarios.

Entities

Institutions

  • arXiv

Sources