Retrieval-Augmented Vision-Language-Action Policies for Cross-Embodiment Generalization in Robotic Manipulation
DOI:
https://doi.org/10.70088/4m6kz628Keywords:
robotic manipulation, vision-language-action policy, retrieval-augmented learning, cross-embodiment generalization, embodiment-aware retrievalAbstract
Robotic manipulation is gradually evolving from performing tasks in fixed scenarios to the flexible deployment of robots, tools, and complex environments. The Vision-Language-Action (VLA) paradigm enables robots to generate executable actions by integrating visual perception with natural language instructions. However, when robots differ in kinematic structures, end-effector configurations, observation perspectives, or action scales, the transferability of learned policies remains severely limited. Existing VLA methods and retrieval-based strategies typically reuse historical task experiences without considering whether those experiences are physically compatible with the current robot embodiment. To address this challenge, this study proposes an enhanced retrieval-augmented VLA framework that explicitly incorporates robot body information into the experience retrieval process. The proposed method improves the applicability of historical trajectories through compatibility screening and target robot action-space mapping. Specifically, the framework retrieves similar operation demonstrations by jointly integrating visual, linguistic, and embodiment features, then selects the most effective experiences based on the physical characteristics of the target robot to generate corresponding actions. Five experiments were conducted on a subset of the publicly available Open X-Embodiment dataset. Results demonstrate that, compared with the strongest baseline, the proposed method achieves notable improvements in action prediction accuracy, cross-embodiment transfer capability, and vision-language retrieval compatibility. These findings indicate that retrieval methods that account for a robot's own physical characteristics can effectively enhance action prediction accuracy, task transferability, and result interpretability in cross-embodiment robotic operations.Downloads
Published
2026-08-01