Role-Stratified Conformal Risk Control Enhances LLM Tool Call Safety
A recent paper on arXiv (2607.24343) presents a novel calibration technique called role-stratified per-field conformal risk control, aimed at enhancing the safety of tool calls made by language-model agents. The authors contend that current methods of conformal risk control evaluate tool calls collectively, potentially masking failures in infrequent high-risk areas. Their strategy involves allocating distinct thresholds and risk budgets to each semantic argument role, ensuring finite-sample assurances for adequately sampled roles while pooling the least common ones. The findings indicate that aggregate certification amplifies the effective budget of a rare role in relation to its occurrence, while role-stratified calibration certifies each role independently. This approach tackles the issue of untrusted content in email bodies, excluding recipient settings, accounts, commands, or credentials. The method functions as a calibration layer, applicable to any per-field detector. The paper was announced as a replace-cross type and is accessible on arXiv.
Key facts
- Paper arXiv:2607.24343 introduces role-stratified per-field conformal risk control.
- Method assigns separate thresholds and risk budgets to each semantic argument role.
- Existing conformal risk control methods certify tool calls as a whole, averaging away rare high-risk field failures.
- Role-stratified calibration certifies each sufficiently sampled role with a finite-sample guarantee.
- Rarest roles are pooled together.
- Untrusted content may shape email bodies but should never set recipient, account, command, or credential.
- Method wraps any per-field detector.
- Announce type is replace-cross.
Entities
Institutions
- arXiv