Observatory index online 41.8781° N / 87.6298° W · Autonomous systems field index How we verify data →
A grid of RT-2 demonstrations pairs natural-language prompts with a robot arm manipulating food, toys, cans, and symbolic targets.
Image: RT-2 project authors / Google DeepMind ↗
SCHOLARLY RECORD / ARXIV

RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Anthony Brohan, Noah Brown, Justice Carbajal, et al.

ConfirmedJul 2023Observatory review: Aug 20, 2026
Educational abstract, not a reproduction

This page preserves only an Observatory-written summary and context. The original source remains authoritative for the paper’s language, methods, findings and limitations.

Plain-language summary

The paper describes vision-language-action models that represent robot actions alongside visual and language inputs.

Why it matters

It is an influential example of adapting large multimodal models to robotic control tasks.

Institutions

  • Google DeepMind

Topics

Observatory directory

Trace any path through the machine world.

Every profile connects to the companies, technologies, industries, research and news around it.

Machines

DronesRobotsHumanoidsAutonomous VehiclesMaritime SystemsSpace Robotics

Ecosystem

CompaniesTechnologiesComponentsIndustriesApplicationsPeople

Intelligence

DefenseResearchLabs & UniversitiesInvestingMarket IntelligenceRegulationNotizie

Explore

FormazioneGlossaryCompareRankingsMappa mondialeTimeline