ARTFEED — Contemporary Art Intelligence

Policy-Invariant Reward Shaping from LLM Feedback: A Framework for Hybrid RL Agents

ai-technology · 2026-08-19

The arXiv preprint 2608.18008 introduces a novel framework for hybrid reinforcement learning agents that merges a large language model (LLM) planner with a reinforcement learning (RL) controller, framed as a Goal-Augmented Markov Decision Process (MDP). The LLM produces a progress score for each state, acting as a bounded potential function, which maintains the integrity of the optimal policy set despite potential inaccuracies in the LLM's scores. This reliability is crucial for practical applications. The study includes numerical validation on a small MDP, exploring four configurations of the potential function and demonstrating robustness even in extreme scaling scenarios. This preprint, accessible on arXiv, enhances research on the integration of LLMs with RL, pertinent to fields like robotics, dialogue systems, and autonomous decision-making.

Key facts

  • The paper formalizes a hybrid LLM-planner and RL-controller architecture as a Goal-Augmented Markov Decision Process.
  • The LLM's per-state progress score is used as a bounded potential function for reward shaping.
  • The shaping term preserves the optimal policy set even when LLM scores are inaccurate.
  • This guarantee is stronger than that of general LLM-as-reward approaches.
  • The result is verified numerically on a small MDP.
  • Four potential configurations are tested, including an adversarial one scaled to twenty times the base reward magnitude.
  • The paper is available as arXiv preprint 2608.18008.
  • The submission is listed under Computer Science > Machine Learning.

Entities

Institutions

  • arXiv
  • arXivLabs
  • Semantic Scholar

Sources