Reinforcement Learning with Symbolic Heuristics for Temporal Planning
A recent paper on arXiv (2505.13372) introduces an advanced Reinforcement Learning (RL) framework designed for generating heuristic guidance in temporal planning. This method emphasizes the use of symbolic heuristics throughout both the RL and planning stages. The researchers define various reward structures and apply symbolic heuristics to address challenges arising from episode truncation in potentially infinite MDPs. Additionally, they suggest learning a residual of an existing symbolic heuristic as a means of correction, rather than developing the entire heuristic anew. This study builds upon earlier work that utilized RL to derive heuristics from the value functions of MDPs formed over training scenarios, with the goal of enhancing planner efficiency in fixed domains. The paper is accessible on arXiv and was released as a replacement type.
Key facts
- Paper arXiv:2505.13372
- Announce Type: replace
- Focus: exploiting symbolic heuristics in RL and planning
- Formalizes different reward schemata
- Uses symbolic heuristics to mitigate episode truncation problems
- Proposes learning a residual of an existing symbolic heuristic
- Builds on prior work using RL for heuristic synthesis
- Aims to improve temporal planner performance
Entities
Institutions
- arXiv