ARTFEED — Contemporary Art Intelligence

Vero: Benchmark for Verified AI Code Generation at Repository Level

ai-technology · 2026-08-15

A new benchmark named Vero has been launched to assess AI agents' proficiency in producing both implementations and machine-verified proofs of their specifications throughout entire software repositories. This benchmark, outlined in a paper on arXiv (2608.13522), fills a void in current verification benchmarks that usually concentrate on single functions or proof generation with existing implementations. Vero includes 43 multi-module instances derived from actual repositories in Python, Dafny, Verus, and Coq, spanning areas such as cryptographic protocols and distributed systems. Each instance challenges the agent to make consistent implementation and proof decisions across various modules, evaluating its ability to manage intricate, real-world codebases. This benchmark represents the first effort to assess combined implementation and proof synthesis at the repository level, with the goal of promoting reliable AI-generated software.

Key facts

  • Vero is the first benchmark for repository-level verified code generation.
  • It contains 43 multi-module instances from real-world repositories.
  • Instances span Python, Dafny, Verus, and Coq.
  • Domains include cryptographic protocols and distributed systems.
  • The benchmark evaluates both implementation and proof synthesis.
  • It addresses the open question of coherent choices across modules.
  • The paper is available on arXiv with ID 2608.13522.
  • The announcement type is cross.

Entities

Institutions

  • arXiv

Sources