ARTFEED — Contemporary Art Intelligence

Role-Stratified Conformal Risk Control Enhances LLM Tool Call Safety

ai-technology · 2026-08-03

A recent paper on arXiv (2607.24343) presents a novel calibration technique called role-stratified per-field conformal risk control, aimed at enhancing the safety of tool calls made by language-model agents. The authors contend that current methods of conformal risk control evaluate tool calls collectively, potentially masking failures in infrequent high-risk areas. Their strategy involves allocating distinct thresholds and risk budgets to each semantic argument role, ensuring finite-sample assurances for adequately sampled roles while pooling the least common ones. The findings indicate that aggregate certification amplifies the effective budget of a rare role in relation to its occurrence, while role-stratified calibration certifies each role independently. This approach tackles the issue of untrusted content in email bodies, excluding recipient settings, accounts, commands, or credentials. The method functions as a calibration layer, applicable to any per-field detector. The paper was announced as a replace-cross type and is accessible on arXiv.

Key facts

  • Paper arXiv:2607.24343 introduces role-stratified per-field conformal risk control.
  • Method assigns separate thresholds and risk budgets to each semantic argument role.
  • Existing conformal risk control methods certify tool calls as a whole, averaging away rare high-risk field failures.
  • Role-stratified calibration certifies each sufficiently sampled role with a finite-sample guarantee.
  • Rarest roles are pooled together.
  • Untrusted content may shape email bodies but should never set recipient, account, command, or credential.
  • Method wraps any per-field detector.
  • Announce type is replace-cross.

Entities

Institutions

  • arXiv

Sources