CoTinyVLA: Sub-Billion-Parameter Robot Model Achieves 90.8% on LIBERO-Plus
Researchers have developed CoTinyVLA, a 0.9-billion-parameter Vision-Language-Action model that achieves 90.8% accuracy on the LIBERO-Plus robustness benchmark. The model uses a Qwen3.5-0.8B backbone and employs chain-of-thought distillation from a 35-billion-parameter teacher model. It processes dual-view temporal input with 16 history frames per step and uses hierarchical CoT spans for planning and reasoning. Paraphrase augmentation expands 40 base commands into 800 variants. LIBERO-Plus includes 10,030 perturbed tasks across seven perturbation dimensions. The work addresses memory constraints in embedded robotics by structuring supervision rather than enlarging the model.
Key facts
- CoTinyVLA has 0.9 billion parameters
- Uses Qwen3.5-0.8B backbone
- Achieves 90.8% on LIBERO-Plus benchmark
- Distills chain-of-thought from 35B teacher model
- Processes 16 history frames per step with dual-view temporal input
- Includes hierarchical Plan and Think spans
- Paraphrase augmentation expands 40 commands to 800 variants
- LIBERO-Plus covers 10,030 perturbed tasks across 7 dimensions
Entities
Institutions
- arXiv