TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models
A new paper on arXiv (2608.08491) introduces TrustRoboReward, a multi-paradigm reward modeling framework designed to address inconsistencies in reward models for reinforcement learning in embodied AI. The framework is equipped with Preference-Ordered Isotonic Score Editing (POISE), which harmonizes pairwise preferences and pointwise scores. The authors construct a unified four-paradigm dataset that includes trajectory progress scoring, video-QA annotations, and other supervision signals. This approach aims to improve long-horizon robotic manipulation by providing scalable vision feedback beyond handcrafted rewards or task-specific annotations. The paper highlights limitations in existing open-source VLM reward judges like RoboReward, which use simple 1–5 trajectory progress scoring and lack pairwise preferences for RLHF, DPO, and Bradley-Terry frameworks. The proposed method also addresses issues with aggregation methods such as TrustJudge, which cannot resolve inconsistencies between pairwise preferences and pointwise scores. The work is relevant to the fields of AI, robotics, and reinforcement learning, and is published as a preprint on arXiv.
Key facts
- Paper arXiv:2608.08491 introduces TrustRoboReward, a multi-paradigm reward modeling framework.
- TrustRoboReward uses Preference-Ordered Isotonic Score Editing (POISE) to harmonize pairwise preferences and pointwise scores.
- The framework constructs a unified four-paradigm dataset including trajectory progress scoring and video-QA annotations.
- Existing open-source VLM reward judges like RoboReward use simple 1–5 trajectory progress scoring and lack pairwise preferences.
- The paper addresses limitations of aggregation methods such as TrustJudge.
- The work targets long-horizon robotic manipulation in embodied AI.
- The paper is a preprint on arXiv with announcement type new.
- The approach aims to provide scalable vision feedback beyond handcrafted rewards.
Entities
Institutions
- arXiv