ARTFEED — Contemporary Art Intelligence

AMD ROCm Enables CUDA-Free Vision-Language-Action Robot Manipulation Pipeline

ai-technology · 2026-07-29

A new study on arXiv (2607.22997) presents a complete technology framework for embodied manipulation, specifically tailored for AMD hardware. It demonstrates that vision-language-action (VLA) policies can be effectively trained outside of a CUDA environment. The setup uses AMD's data-center training chips, Radeon PRO GPUs for both simulation and rendering, along with Ryzen AI for edge computing, all facilitated by the open ROCm software platform. The research includes four demonstrations: (1) a Sim-to-Real manipulation pipeline featuring SmolVLA on a physical Franka arm, and (2) a language-directed semantic manipulation task. This aligns with the Physical AI trend, echoing statements from industry leaders like Jensen Huang and Dr. Lisa Su, who both emphasize the significance of AI in the physical realm.

Key facts

  • arXiv paper 2607.22997 presents an AMD-accelerated pipeline for VLA-based robot manipulation.
  • The pipeline uses AMD data-center silicon, Radeon PRO GPUs, and Ryzen AI edge compute.
  • All components are unified by the open ROCm software stack.
  • Demonstrates that CUDA is not required for training and deploying VLA manipulation policies.
  • Includes four progressive demonstrations: Sim-to-Real with SmolVLA on a Franka arm, and semantic language-guided manipulation.
  • Jensen Huang stated Physical AI is 'the next big thing' at GTC Paris, June 2025.
  • Dr. Lisa Su said 'we're entering the world of Physical AI' at CES 2026.
  • The paper is categorized as an arXiv cross submission.

Entities

Institutions

  • AMD
  • ROCm
  • Franka
  • SmolVLA
  • arXiv

Locations

  • Paris
  • France

Sources