Surgical WAM: Action-Free Video Pretraining Boosts Surgical Robot Learning
A new research paper introduces the Surgical World-Action Model (Surgical WAM), a unified generative model designed to improve surgical robot learning by leveraging action-free endoscopic video pretraining. The study addresses the scarcity of action-labeled demonstrations in surgical robotics, which are costly to collect due to the need for teleoperated trajectories with synchronized kinematics. The authors propose that abundant endoscopic video can be used to learn world models of surgical scenes, but existing models rarely translate learned dynamics into closed-loop control. Surgical WAM aims to answer whether action-free video pretraining improves closed-loop surgical manipulation under a fixed budget of action-labeled demonstrations. The paper is available on arXiv under the identifier 2608.11204 and was announced as a cross-type submission. The research highlights a potential shift toward data-efficient learning in surgical robotics, emphasizing the use of readily available video data to enhance policy learning without extensive human annotation.
Key facts
- Surgical WAM is a unified generative model for surgical robot learning.
- It uses action-free endoscopic video pretraining to address data scarcity.
- The model aims to improve closed-loop surgical manipulation.
- Existing surgical world models rarely translate learned dynamics into control.
- The paper is available on arXiv with ID 2608.11204.
- The research focuses on data-efficient learning for surgical robots.
- Teleoperated trajectories with synchronized kinematics are costly to collect.
- Endoscopic video is relatively inexpensive and abundant.
Entities
Institutions
- arXiv