arXiv paper presents RGE, an ontology-based trust monitor for long-horizon AI agents
The paper arXiv:2608.17718v1 presents a novel approach for monitoring long-horizon AI agents by evaluating their changing trajectories in relation to user-approved tasks. It emphasizes the importance of considering the entire trajectory rather than just isolated actions, as unnoticed drift can occur. Existing methods typically involve local compliance evaluations, final assessments, or broad risk scoring, which do not include prefix-level analysis. The concept of 'ontological trust' is defined as a property of trajectory prefixes conditioned by tasks and is implemented in RGE, an online monitoring system that breaks down trust into Role, Goal, and Evidence. RGE employs large language models for organized task representations, while other systems handle trust-state updates and interventions. This research aims to enhance AI alignment and improve early drift detection compared to current techniques.
Key facts
- The paper introduces 'ontological trust,' a task-conditioned property of trajectory prefixes.
- RGE is an online monitor that decomposes trust into Role, Goal, and Evidence.
- Long-horizon agents operate across many steps, tools, and observations.
- Existing monitors mostly check local compliance, deliver final-trace verdicts, or score generic risk.
- Drift can occur when an agent calls the right tool with plausible arguments but shifts toward a broader role or adjacent objective.
- RGE uses LLMs only to derive structured task and step representations.
- The paper is catalogued as arXiv:2608.17718v1 and announced as a new submission.
- No author names, institutions, or locations are named in the abstract.
Entities
Institutions
- arXiv