DreamGuard: Proactive Runtime Guardrail for LLM Agents via Risk-Aware World Model
A new research paper on arXiv (2608.05695) introduces DreamGuard, a proactive runtime guardrail for large language model (LLM) agents. The system uses a risk-aware world model to predict future latent states, enabling it to detect both immediate hazards and prefix risks that could lead to long-horizon dangers. Unlike reactive guardrails that only assess the current action, DreamGuard models how risk evolves across the trajectory, addressing a critical blind spot for actions that appear benign individually but can drift agents toward hazardous states. The paper was announced as a new submission and is available at the provided URL.
Key facts
- Paper ID: arXiv:2608.05695
- Announcement type: new
- Proposes DreamGuard, a proactive guardrail for LLM agents
- Uses a risk-aware world model with a compact recurrent latent state
- Predicts future latent states to derive immediate-hazard and prefix-risk evidence
- Addresses long-horizon risks where benign-looking actions can lead to hazardous states
- Focuses on preventing unsafe actions when LLM agents interact with external tools and real-world systems
- Available at https://arxiv.org/abs/2608.05695
Entities
Institutions
- arXiv