CompressAgent: Benchmarking Reliability of Compressed Tool-Using Agent Controls
A recent study introduces CompressAgent, a benchmark designed to assess the reliability of compressed agent control contexts (ACCs) within tool-using language models, now accessible on arXiv (2608.01056). This research fills existing gaps in evaluating operational reliability concerning prompt compression. CompressAgent evaluates nine ACCs across three distinct task categories, employing three Qwen API model identifiers, six context budgets, and a total of 15,525 runs. Findings reveal a nonlinear reliability frontier; with 75% of context retained, generic rewriting and section-based compression achieve success rates of 92.7% and 92.4%, respectively, closely approaching the 93.8% baseline for full context. The benchmark is verified within its environment, emphasizing the balance between compression and reliability, which is vital for developers using compressed models in practical scenarios.
Key facts
- CompressAgent is a benchmark for ACC compression reliability.
- It covers nine ACCs, three task families, three Qwen API models, six budgets, and 15,525 runs.
- At 75% retained context, generic rewriting achieves 92.7% success.
- Section-based compression achieves 92.4% success at 75% retained context.
- Full-context baseline success rate is 93.8%.
- Between 50% and 35% retained context, methods diverge sharply.
- At 35% retained context, performance drops significantly for several methods.
- The benchmark is environment-verified.
- The paper is available on arXiv with ID 2608.01056.
Entities
Institutions
- arXiv
- Qwen