Image: RT-2 project authors / Google DeepMind01vision-language-action modelsRT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
Anthony Brohan, Noah Brown, Justice Carbajal, et al.
The paper describes vision-language-action models that represent robot actions alongside visual and language inputs.