ARTFEED — Contemporary Art Intelligence

CoPES: Cooperative Coevolution for Efficient LLM Agent Post-Training

ai-technology · 2026-08-04

A novel approach known as Cooperative Parameter-subspace Evolution Strategy (CoPES) has been developed to tackle the substantial GPU-hour demands of evolution strategies (ES) in post-training large language model (LLM) agents, especially in resource-limited situations. This technique breaks down the entire parameter space into smaller, manageable subspaces and collaborates on them to enhance optimization efficiency. The research involved post-training a Qwen3.5-4B tool-using agent for mathematical tasks, testing it across five benchmarks of differing difficulty levels. The paper can be found on arXiv under ID 2608.02391. CoPES aims to achieve performance comparable to gradient-based reinforcement learning (RL) while being memory-efficient, as ES allows for comprehensive parameter post-training without backpropagation. In environments with limited GPUs, however, ES's high GPU-hour requirements can lead to excessively lengthy training periods. CoPES seeks to mitigate this issue through cooperative coevolution, which is anticipated to shorten training times while preserving performance. This study holds significance in the realms of AI and machine learning, particularly for enhancing LLM agents in constrained settings.

Key facts

  • CoPES is a cooperative coevolutionary method for post-training LLM agents.
  • It decomposes the full parameter space into lower-dimensional subspaces.
  • The method is designed for resource-constrained settings with few GPUs.
  • Post-training is performed on a Qwen3.5-4B tool-using agent for math tasks.
  • Evaluation is done on five benchmarks of varying difficulty.
  • The paper is available on arXiv with ID 2608.02391.
  • ES enables memory-efficient full-parameter post-training without backpropagation.
  • CoPES aims to match gradient-based RL performance while reducing GPU-hour requirements.

Entities

Institutions

  • arXiv

Sources