Causal Audit Challenges Latent Communication Claims in Multi-Agent LLMs
A recent preprint on arXiv (2608.04893) conducts a causal analysis of relayed key-value (KV) caches within multi-agent large language model (LLM) systems, questioning the belief that sharing KV caches enhances performance. This analysis involves substituting the relayed cache with deranged, zeroed, and moment-matched random alternatives. Two experimental scenarios are explored: one necessitating the sender's private data, where answer-relevant relays achieve 100% performance compared to 23–25% for answer-irrelevant relays, consistent across three model families and five checkpoints. The second scenario employs a five-seed protocol, revealing an equivalence within 2.8 points under Holm-corrected TOST on GSM8K and ARC-Challenge. The results indicate that the advantages of relayed KV caches hinge on the receiver's requirement for private information.
Key facts
- Preprint arXiv:2608.04893 presents a causal audit of relayed KV caches in multi-agent LLMs.
- The audit replaces relayed caches with deranged, zeroed, and moment-matched random counterparts.
- Two regimes are tested: one where the receiver needs the sender's private information, and one where it does not.
- In the information-required regime, answer-relevant relays achieve 100% performance versus 23–25% for irrelevant relays.
- The contrast is replicated across three model families, five checkpoints, and a prose document-QA surface.
- In the information-not-required regime, a pre-registered five-seed protocol establishes equivalence within 2.8 points.
- Equivalence is shown under Holm-corrected TOST on GSM8K and ARC-Challenge across three Qwen3 scales, and on MedQA at 8B.
- One cell on MedQA shows a small detectable difference.
- The study challenges the assumption that relayed caches always convey useful latent thoughts.
- The work provides a causal methodology for auditing communication protocols in multi-agent systems.
Entities
Institutions
- arXiv