RMSWeb: New Training Recipe for Compact Web Agents
A recent paper on arXiv presents RMSWeb, a novel three-part training methodology for compact web agents utilizing Qwen3-VL-Instruct at both 8B and 32B scales. It tackles the difficulties associated with training web agents, such as costly data acquisition and suboptimal reinforcement learning (RL) following supervised fine-tuning (SFT). RMSWeb includes reflection-conditioned retries to enhance collection efficiency and reduce successful trajectory lengths, failure-mode mining to focus offline RL on essential states, and Salvage-DS, which integrates an action-semantic polarized reward, contrast-and-competence-gated dynamic sampling, and an action-level reward framework. The study notes that full trajectory datasets are largely composed of routine states, and group-relative RL on web actions may experience ineffective or misleading updates. These techniques aim to boost training efficiency and agent capabilities. The paper can be found on arXiv with the identifier 2608.00335.
Key facts
- RMSWeb is a three-part recipe for training compact web agents.
- The recipe is applied to Qwen3-VL-Instruct at 8B and 32B scales.
- Reflection-conditioned retries increase collection yield and shorten successful trajectories.
- Failure-mode mining concentrates offline RL on critical states exposed by the SFT policy.
- Salvage-DS combines an action-semantic polarized reward, contrast-and-competence-gated dynamic sampling, and an action-level reward design.
- The paper addresses challenges in data collection and post-SFT reinforcement learning.
- Full trajectory corpora are dominated by routine states.
- Group-relative RL on web actions can yield weak or misleading relative updates.
- The paper is published on arXiv with ID 2608.00335.
Entities
Institutions
- arXiv