Describe the movement first
Write one sentence about the physical action. Name the robot, the object and the destination. Add only what the footage and its source establish. A sentence about moving a cup is more useful than an unsupported claim about general intelligence.
This worksheet is our editorial method. The OpenVLA example below shows how to use information supplied beside a research video. It does not estimate the reliability of the robot.
Keep five fields beside the clip
| Field | Record |
|---|---|
| Identity | Robot model, version, original uploader and recording date when known |
| Task | Object, starting position, target and visible completion condition |
| Control | Remote control, autonomous operation or an unknown mode |
| Attempts and timing | Full attempts, edits, resets, visible failures and any stated playback speed |
| Conditions | Surface, lighting, supports and people assisting |
Give unknowns a place in the caption
A hand outside the frame may reset the object. A cut can remove a recovery attempt. The absence of a visible controller does not establish autonomy. Use wording such as control mode not stated when the source does not explain it.
Keep observed movement separate from the uploader’s explanation. Cite the original recording and any technical report. Avoid turning a percentage into a reliability claim without the number of trials and a definition of success.
Before writing the post
- Check that the model in the clip matches the specification sheet.
- Keep simulation and physical-robot footage separate.
- State a limitation that changes the meaning of the result.
- Link the original source and correct the page when better evidence arrives.
The original OpenVLA project labels its sample policy clips as playing at 1.5 times normal speed. Other comparison clips on the same page use different speeds. A caption should carry the rate attached to its own clip. Eight seconds of playback at 1.5 times speed represents twelve seconds of recorded movement. This is a timing calculation, not a measured task result. [1]
Sources and verification
- OpenVLA – An Open-Source Vision-Language-Action Model ↗OpenVLA research team · Read 8 October 2026
Original 7B model. Separates direct evaluation from robot-specific fine-tuning.
Article history
Added an OpenVLA playback-speed example and a calculation separating video duration from recorded movement.
Report a correction