Visual Prompting and ACT Enhance Robotic Pick-and-Place in Retail
Robotic pick-and-place in convenience stores often fails due to dense object arrangements, occlusions, and variations in color, shape, size, and texture. A new paper introduces a perception-action pipeline that uses annotation-guided visual prompting: bounding box annotations mark both pickable objects and placement targets, providing structured spatial guidance. Instead of step-by-step planning, the system applies Action Chunking with Transformers (ACT), an imitation learning algorithm that predicts chunked action sequences from human demonstrations, enabling smooth, adaptive operations. Evaluated on success rate and visual grasp analysis, the approach shows improved grasp accuracy and adaptability in retail settings.
Key facts
- Robotic pick-and-place in convenience stores faces challenges from dense arrangements, occlusions, and object variations.
- The paper proposes a perception-action pipeline using annotation-guided visual prompting.
- Bounding box annotations identify pickable objects and placement locations.
- Action Chunking with Transformers (ACT) is used as an imitation learning algorithm.
- ACT predicts chunked action sequences from human demonstrations.
- The system enables smooth, adaptive, data-driven pick-and-place operations.
- Evaluation is based on success rate and visual analysis of grasping behavior.
- Results demonstrate improved grasp accuracy and adaptability in retail environments.
Entities
Institutions
- arXiv
- arXivLabs
- Semantic Scholar