What Odyssey has released
Odyssey introduced Odyssey-3 on 15 September and released its research preview on 8 October 2026. Its robot examples include pouring cereal and closing a screwbox. The central proposal is to reuse a model trained on visual experience, then learn the controls of a particular machine. [1] [2]
A world model predicts how a scene may develop. A robot policy selects an action from observations and a goal. Connecting the two requires a learned mapping from scene information to commands that the hardware can execute. Watching a plausible generated movement cannot establish that this mapping works on a physical arm.
From visual experience to a control interface
Odyssey calls the model an autoregressive diffusion transformer. Training mixes annotated internet video, gameplay with controls and simulated rigid-body interactions. The release describes distillation for interactive generation. It does not supply a complete robotics architecture or reproducible training recipe. [2]
Autoregressive prediction uses earlier context to produce the next output. Diffusion generates a sample through repeated denoising. These terms identify model families; the robot still needs a defined observation and command interface.
A learned visual representation is the model’s numerical encoding of its observations. It can carry information useful for predicting motion without exposing named coordinates or a readable list of physical laws. An action decoder is a trained output component that maps these representations to a machine’s controls. Odyssey trains that component using paired observations and actions. [1]
For a new gripper, the developer must define what an action means. A position target, a joint target and a motor torque have different units and consequences. Camera timestamps must also match the actions they describe. These are interface requirements for a working controller; the announcement does not establish which command format its manipulation policies use.
Following a cereal-pouring task
The cereal-pouring example gives a useful way to explain feedback. The sequence below is a conceptual reading of the task, not a reconstruction of unpublished Odyssey software. It makes no claim about camera placement, force sensing, action frequency or internal planning.
| Stage | Information needed by the controller |
|---|---|
| Observe | Acquire the scene containing the container, bowl and gripper. The current relationship between them matters after every movement. |
| Represent | Encode the observations into features a policy can use. A named bowl coordinate or explicit physics solver cannot be assumed. |
| Adapt | Learn which robot commands correspond to the recorded observations and task examples. The hardware interface defines the output units. |
| Act | Send the chosen command to the robot controller. That controller must handle the permitted motion and actuator limits. |
| Observe again | Read the resulting scene. The container may have moved, remained on the table or slipped. |
| Correct | Use fresh feedback to choose another action. A missed grasp requires a different response from a successful lift. |
Closed-loop control uses the consequences of an action to choose what happens next. During pouring, a fixed wrist motion can produce different outcomes with different openings or fill levels. A useful evaluation would therefore record spills, amount transferred and recovery attempts, alongside completion. These are proposed checks, not results published for Odyssey-3.
What the physical robot examples establish
Odyssey reports adaptation from tens of hours of arm training recordings and recovery actions absent from those recordings, including gripper reorientation and retrieval of a dropped object. Flexion built humanoid policies using tens of hours of teleoperation. Listed tasks include opening a blue container to retrieve a cardboard box, and placing a mug on a plate. [1]
Flexion independently confirms the collaboration through its company account. Odyssey says the policies handled environmental and lighting changes better than the VLA baselines it tested. VLA means vision-language-action, a policy that connects visual input and language to actions. The release provides no baseline names, trial counts or success percentages for this comparison. [3] [2]
Flexion’s earlier Reflect description separates mission planning, motion policies, whole-body control and runtime services. That June system is useful context for the work around a humanoid policy. Its reported metrics cannot be transferred to Odyssey-3, and its complete architecture cannot be assumed to describe the later collaboration. [4]
The recovery examples support a narrow observation that selected runs continued after a mistake. Measuring reliability requires counting unsuccessful recoveries too. The relevant denominator is every attempted task under a stated setup, including runs stopped by an operator.
What the physics score measures
Physics-IQ starts from 66 physical experiments recorded from three viewpoints and repeated twice. Its video-to-video task supplies three seconds of context and evaluates a five-second continuation. Metrics compare the location, timing and strength of visual changes, plus pixel error. The Verified revision corrects prompts and reference artifacts and gives samples and metrics equal weight. [5]
| Odyssey-3 Pro condition | Video-to-video | Image-to-video |
|---|---|---|
| Standard chart point | 63.4 | 50.0 |
| Best-of-eight selection | 66.1 | 54.7 |
The chart reports four-run averages for standard results and one run for best-of-eight. Selection from eight candidates changes the sampling budget. Neither score measures a robot gripping an object, regulating contact force or recovering balance. [2]
The benchmark has a recorded physical reference, so it asks a more specific question than whether a clip looks convincing. Even a close video continuation leaves the robot-control question open. The policy still needs to execute commands under sensor error, contact and timing constraints.
Generated worlds have a separate evaluation
WorldMark tests action following, scene memory and image quality over 500 cases. Its 100 starting images each receive five action sequences. Sequences last 20, 40 or 60 seconds. A return movement checks whether the earlier scene survives; another metric checks response delay after a command changes. [6]
Odyssey reports first place in three of four WorldMark splits using the mean of 13 reported scores. That aggregation includes appearance and memory alongside motion. It cannot be read as a measured gain in robot manipulation. [2]
For a reader comparing systems, keep three records separate. Generated-world tests measure the simulated scene. Physics prediction compares generated frames with recordings. Robot trials measure actions on hardware. A result in one record supplies no missing trial count for another.
What a new task would need
For another manipulation task, start with synchronized observations and commands from the intended robot. Vary object poses and camera conditions, and keep held-out objects or arrangements for testing. Record successful actions, failed attempts and operator interventions. This is an evaluation proposal, not a claim about Odyssey’s undisclosed dataset.
- Specify hardware, cameras, command units and the software version for each result.
- Report the number of training episodes and held-out trials, beyond a total duration in hours.
- Define task completion, allowed retries, time limits and intervention rules before testing.
- Measure observation-to-command latency and its variation while the robot is moving.
- Compare against named baselines under the same data budget and physical conditions.
The published material leaves those robotics details unresolved, along with exact dataset size, training compute, failure rates and the boundary of generalization. A useful next result would report every attempt for one named robot and task, with the changed conditions and recovery outcomes visible.
Sources and verification
- Odyssey-3 introduction and adaptation experiments ↗Odyssey · Read 10 October 2026
Dated 15 September 2026. Company account of robot experiments and action-decoder training; no complete robotics evaluation protocol supplied.
- Odyssey-3 release and evaluation charts ↗Odyssey · Read 10 October 2026
Release dated 8 October 2026. Physics-IQ chart snapshot dated 7 October. Best-of-eight labels checked in the original chart markup. Company evaluations.
- Flexion confirms its Odyssey collaboration ↗Flexion official company account · Read 10 October 2026
Company post links the September introduction. Partnership confirmation, without a trial table.
- Flexion Reflect v1.0 system description ↗Flexion · Read 10 October 2026
Dated 29 June 2026. Earlier system context, not a specification or benchmark for the Odyssey collaboration.
- Physics-IQ Verified paper ↗Tim Rädsch and colleagues, arXiv · Read 10 October 2026
Version 1, 17 June 2026. Sections 2 and 3 define the experiment clips, metrics and scoring corrections. It does not independently evaluate Odyssey-3.
- WorldMark benchmark design and results ↗Alaya Lab and collaborators · Read 10 October 2026
Benchmark page describes 500 cases and nine metrics. Four motion metrics each have translation and rotation columns, producing 13 reported scores per split.
Article history
Prepared a source-checked explanation separating robot trials, physical prediction and generated-world evaluation.
Report a correction