CloudEdgeVLA: New Framework for Cloud-Edge Robot Control
The newly introduced CloudEdgeVLA framework tackles the issue of implementing billion-parameter Vision-Language-Action (VLA) models in mobile robots, which must balance the necessity for cloud-based semantic reasoning with the demand for low-latency local control. This innovative system views temporal misalignment as a challenge in representation learning, where a cloud VLA processes delayed observations into gradually changing task features. Simultaneously, a streamlined edge head merges the latest cloud features with real-time local vision. For training, the model associates current frames with randomly delayed ones, targeting the same action in both fresh and outdated paths, promoting the retention of task-level information in the cloud representation while the edge path provides prompt reactions. The full paper can be found on arXiv with the identifier 2608.00569.
Key facts
- CloudEdgeVLA is a cloud-edge policy for VLA models.
- It addresses network delay and jitter in mobile robot control.
- The cloud VLA encodes delayed observations into task features.
- A lightweight edge head combines cloud features with local vision.
- Training pairs current and delayed frames with the same action target.
- The approach treats temporal misalignment as a representation-learning problem.
- The paper is on arXiv with ID 2608.00569.
Entities
Institutions
- arXiv