
SCHOLARLY RECORD / ARXIV
RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
Anthony Brohan, Noah Brown, Justice Carbajal, et al.
Educational abstract, not a reproduction
This page preserves only an Observatory-written summary and context. The original source remains authoritative for the paper’s language, methods, findings and limitations.
Plain-language summary
The paper describes vision-language-action models that represent robot actions alongside visual and language inputs.
Why it matters
It is an influential example of adapting large multimodal models to robotic control tasks.
Institutions
- Google DeepMind