Argus: A Persistent Agentic Runtime for Long-Horizon Reasoning
A recent paper on arXiv (ID: 2608.05144) presents Argus, a self-evolving runtime tailored for long-horizon reasoning tasks. This innovative system features a multi-agent framework that includes roles such as Manager, Planner, Engineer, and Reviewer, all tasked with executing bounded missions on a stable project state. Argus distinguishes between user intent and operational goals, constraints, and verification standards, incorporating memories, skills, procedures, and routing decisions only after thorough reviews and task-specific verifications. While model weights remain unchanged, self-evolution is driven by persistent runtime states and control policies, allowing for autonomous actions between operator-defined escalation points. In tests across seven GPT-5.5 environments, Argus scored about 78% on SWE-Bench Pro, outperforming Direct Copilot's 59%, using 1.41 times the total resources. The paper is newly submitted and can be accessed via the arXiv link.
Key facts
- Argus is a persistent, self-evolving runtime for long-horizon reasoning.
- It uses Manager, Planner, Engineer, and Reviewer agents.
- It separates stable user intent from operational objectives and constraints.
- Self-evolution occurs through persistent runtime state and control policy.
- Model weights remain fixed.
- Argus achieves about 78% on SWE-Bench Pro vs 59% for Direct Copilot.
- It uses 1.41 times the aggregate resources.
- The paper is on arXiv with ID 2608.05144.
Entities
Institutions
- arXiv