MDArena Benchmark Evaluates Coding Agents on Molecular Dynamics Tasks
The introduction of MDArena marks a new standard for assessing the effectiveness of coding agents in realistic molecular dynamics (MD) workflows. This benchmark features 50 containerized tasks sourced from active biomolecular simulation projects, encompassing 29 molecular systems and 14 diverse research protocols, such as trajectory analysis, preparation of complex systems, free-energy protocols, and enhanced sampling. Six model/harness configurations were tested, including Codex and OpenCode. Notably, Codex GPT-5.5, when utilizing extra-high reasoning effort, excelled with 24 out of 50 Strict-Pass@1 successes (48%). It was followed by Codex GPT-5.5 Medium with 21 successes and OpenCode Gem. This research underscores the ability of coding agents to significantly streamline computational biomolecular simulation workflows, thus fostering scientific advancement. The study addresses the gap in evaluating coding agent reliability for realistic MD tasks, and the findings are published on arXiv with the identifier 2608.02642.
Key facts
- MDArena is a benchmark of 50 containerized tasks from active biomolecular simulation projects.
- The benchmark spans 29 molecular systems and 14 broad research protocols.
- Protocols include trajectory analysis, complex system preparation, free-energy protocols, and enhanced sampling.
- Six model/harness configurations were evaluated, including Codex and OpenCode.
- Codex GPT-5.5 at extra-high reasoning effort achieved 24/50 Strict-Pass@1 successes (48%).
- Codex GPT-5.5 Medium achieved 21/50 successes.
- OpenCode Gem also performed, but exact score not specified in the abstract.
- The paper is available on arXiv (2608.02642).
Entities
Institutions
- arXiv