Treatment Leakage Found in MCP Agent Security Evaluation
An assessment of security concerning tool-using agents revealed treatment leakage, where the labels stored do not accurately represent behavioral realities. This audit, which relied on a preserved campaign, traced 10,200 execution rows to 180 requests bound to models, 45 semantic requests, and 15 observable stimuli. Although two schema treatments were implemented, the intended external payload-family corpus was absent. The historical grader showed clear treatment leakage: treatment metadata restricted the ATTACK_SUCCESS class, indicating that fixed behavior could alter class under treatment relabeling. A treatment-blind reconstruction amended 58 historical ATTACK_SUCCESS or HIJACK_ATTEMPT labels to authorized benign completions, while maintaining three verified protected-data transfers and one separate unauthorized-forwarding incident. The locked v2 census shows no ATTACK_SUCCESS records, with the forwarding incident still classified as a HIJACK_ATTEMPT at a semantic boundary related to objective completion. A dual-reviewer blinded concordance review was performed, emphasizing the significance of construct validity in security evaluations of AI agents.
Key facts
- 10,200 execution rows were traced to 180 model-bound requests, 45 semantic requests, and 15 observable stimuli.
- Two schema treatments were delivered, but the planned external payload-family corpus was not.
- Treatment metadata gated the ATTACK_SUCCESS class, causing direct treatment leakage.
- 58 historical ATTACK_SUCCESS or HIJACK_ATTEMPT labels were corrected to authorized benign completions.
- Three verified protected-data transfers and one separate unauthorized-forwarding case were preserved.
- The locked v2 census contains exactly zero ATTACK_SUCCESS records.
- The forwarding case remains a HIJACK_ATTEMPT at a semantic boundary concerning objective completion.
- A dual-reviewer blinded concordance review was conducted.
Entities
—