ROBOTICS FIELD NOTESENGLISH EDITION / 8 October 2026
Research

Vision-language-action models explained

How RT-2 and OpenVLA turn images and instructions into robot actions, and what their results do not establish.

An action model has an action output

A vision-language-action model uses visual information and a language instruction to predict actions for a robot. RT-2, published by Google DeepMind in July 2023, adapted pretrained vision-language models using robot trajectories alongside web tasks. The resulting model could produce robot action tokens as well as process visual and linguistic concepts. [1]

The word action deserves attention. An answer such as 'pick up the cup' leaves the movement unspecified. In RT-2, the model output describes a step of end-effector motion and the gripper state. It is an interface to robot control, with a defined representation rather than an unrestricted instruction to any machine. [1]

What an RT-2 action contains

ElementMeaning
Continue or terminate flagWhether the current episode should continue
Three position changesRequested end-effector displacement
Three rotation changesRequested orientation adjustment
Gripper valueRequested gripper opening state

[1]

This action representation belongs to the robot setup described in RT-2. It should not be copied onto a humanoid specification as a count of joints or as proof of whole-body walking control. [1]

Web knowledge has a physical boundary

The RT-2 paper reports experiments in which language and visual pretraining helped robots apply existing manipulation skills to new semantic instructions. Its limitations section makes a narrower claim than many headlines. Additional web experience did not by itself give the robot new physical motions. The available motor skills remained bounded by those in the robot training data. [2]

That is a useful distinction when reading a result. Recognizing which item satisfies an instruction and successfully moving it are connected tests, but success at the first does not supply missing reach, grip force or a previously untrained movement. The authors also identify computation cost as a concern for tasks requiring faster control. [2]

What OpenVLA makes inspectable

The original OpenVLA release describes a 7-billion-parameter model trained on 970,000 robot episodes from Open X-Embodiment. Its model weights and training pipeline are available for research. Its architecture combines visual encoders with a language-model backbone that predicts tokenized actions, which are decoded into continuous robot commands. [3]

The project separates direct tests on existing robot setups from fine-tuning on Franka arm tasks. That distinction matters when a result is called generalization. The reported comparison also finds that a task-specific Diffusion Policy performs better on some narrow, precise tasks. A larger general model is not automatically the best controller for each operation. [3]

Later VLA systems can use different action outputs. The 2025 OpenVLA-OFT work combines parallel decoding, action chunks and continuous action values. Its name still describes vision, language and action, even though its decoding differs from original OpenVLA. Check the action representation of the version being discussed. [4]

Four details to check

Check the robot, action format, training overlap and adaptation data. Then read the success definition and failure cases.

DetailWhy it changes the claim
Action spaceA gripper command is not a full-body command
Training overlapAn unseen instruction can involve familiar motions
Adaptation dataA new setup may require further training
Control rate and latencyActions must reach the hardware in time

Sources and verification

  1. RT-2 – New model translates vision and language into action ↗Google DeepMind · Read 8 October 2026

    July 2023 explanation of RT-2 and its discrete action representation.

  2. RT-2 – Vision-Language-Action Models Transfer Web Knowledge to Robotic Control ↗Brohan et al. · Read 8 October 2026

    Section 5 states that web pretraining does not itself add new physical motions.

  3. OpenVLA – An Open-Source Vision-Language-Action Model ↗OpenVLA research team · Read 8 October 2026

    Original 7B model. Separates direct evaluation from robot-specific fine-tuning.

  4. Fine-Tuning Vision-Language-Action Models Optimizing Speed and Success ↗Moo Jin Kim, Chelsea Finn and Percy Liang · Read 8 October 2026

    2025 project. Used for continuous action representation and parallel action chunk decoding, not a current performance ranking.

Article history

Added the 2025 OpenVLA-OFT action representation and decoding example.

Report a correction