Ant Group's Robbyant claims breakthrough in robot vision with LingBot-Vision model
Ant Group's embodied AI unit, Robbyant, claims its new vision model LingBot-Vision surpasses Meta's DINOv3 in object edge recognition, achieving higher precision on the NYUv2 benchmark with fewer parameters and less training data. The model is the first trained specifically to recognize object edges, enabling robots to understand 3D spaces with high precision. LingBot-Vision powers LingBot-Depth 2.0.
Key facts
- Ant Group's Robbyant unit developed LingBot-Vision, a perception model for embodied AI.
- LingBot-Vision surpasses Meta's 7-billion-parameter DINOv3 model on the NYUv2 depth-estimation benchmark.
- LingBot-Vision uses one-seventh as many parameters and less than a third of the training data compared to DINOv3.
- It is the first model trained specifically to recognize object edges.
- The AI can pinpoint boundaries down to a fraction of a single pixel.
- LingBot-Vision serves as the engine for LingBot-Depth 2.0.
- The claims are based on a research paper published by the Robbyant team.
- The model aims to improve robot sensing of 3D spaces.
Entities
Institutions
- Ant Group
- Robbyant
- Meta Platforms