New Benchmark TRACE for Human-AI Controller Coordination Under Drift and Failure
TRACE has been developed by researchers as a multi-layer benchmark to assess the collaboration among human operators, AI decision-making modules, and automated control systems in cyber-physical and AI-enhanced environments. This benchmark tackles the issue of diagnosing drift—variations that may arise from any level of the system stack—by offering time-synchronized, multi-layer traces that illustrate the propagation of drift and failures. The dataset, comprising 1,918 drifted traces, originates from ALFRED, a benchmark focused on grounded instructions for common household activities. Each trace consists of a time-aligned series of records across five execution layers: state, observation, decision, rules, and control, annotated with drift type and onset time. This research seeks to address a gap in current benchmarks, which often lack the necessary detail to pinpoint coordination failures. The paper can be found on arXiv under the identifier 2608.06657.
Key facts
- TRACE is a new benchmark for human-AI controller coordination.
- It focuses on drift and failures in multi-layer systems.
- The dataset is derived from ALFRED, a household task benchmark.
- It includes 1,918 drifted traces.
- Each trace covers five execution layers: state, observation, decision, rules, control.
- Traces are time-aligned and labeled with drift type and onset time.
- The benchmark aims to diagnose where coordination breaks down.
- The paper is available on arXiv (2608.06657).
Entities
Institutions
- arXiv
- ALFRED