ARTFEED — Contemporary Art Intelligence

RATTL Introduces Belief-Dependent Robustness for Safe Sequential Decision-Making

ai-technology · 2026-08-19

The study titled RATTL (Risk-Adversarial Total-Reward Learning), presented in arXiv preprint 2608.17574, investigates the level of caution an agent should adopt while navigating its environment. This caution is connected to epistemic uncertainty through a Bayesian posterior, leading to the creation of a Wasserstein ambiguity set, with its radius being a monotonic function of the posterior. As more evidence is collected, the radius decreases, influencing the agent's planning behavior between maximizing total rewards and ensuring worst-case robustness. The authors introduce a Safety Sandwich result, demonstrating that the RATTL value function is bounded by an uninformed robust value and the optimal solution with complete knowledge, ultimately converging to the best outcome as data is acquired. The approach effectively aligns robustness with the agent's knowledge state while remaining mathematically manageable.

Key facts

  • RATTL stands for Risk-Adversarial Total-Reward Learning.
  • It ties caution to epistemic uncertainty in sequential decision-making.
  • The agent holds a Bayesian posterior over unknown dynamics.
  • A Wasserstein ambiguity set is used, with radius as a monotone function of the posterior.
  • The radius contracts with evidence, adapting behavior between worst-case robustness and risk-neutral maximization.
  • The design follows a duality underlying the Entropic Value-at-Risk.
  • The planning problem is well posed under transience and compactness conditions.
  • A Safety Sandwich theorem shows the value lies between uninformed robust and full-knowledge optimum, with gap vanishing as posterior concentrates.

Entities

Sources