ROBOTICS FIELD NOTESENGLISH EDITION / 10 October 2026
Research / 2026

How Reflex Helps a Unitree G1 Catch Flying Boxes

How Reflex uses RGB-D history and whole-body control to catch boxes, with separate simulation results, physical trials and limits of the evidence.

The box arrives before the robot can start over

Reflex studies a short, unforgiving task. A person throws a carton, and a Unitree G1 must move its body into the flight path and keep the box after contact. The authors report 65% success at 2.5 m and 40% at 5 m on the physical robot. Their 85.7% RGB-D result comes from simulation. [1]

Taoyang Jia, Weikai Huang and Linxin Song share first authorship. The team spans the University of Washington, Stanford, USC, the National University of Singapore, UNC Chapel Hill and the Allen Institute for AI. Their public repository still lists the training code, checkpoints and deployment stack as forthcoming on October 10, 2026. [2]

Two images carry different information from one

The model takes torso-camera RGB-D at 192 × 144 pixels. A convolutional network (CNN) localizes the box; depth supplies a surface point. A recurrent network (GRU) reads 25-frame histories with camera motion and produces an eight-value latent. Robot/reference state and this latent drive 29 joint-position targets at 50 Hz. [3]

RGB identifies a location in the image. Depth places that visible surface along the camera’s viewing ray. With calibrated focal length fx and image center cx, the horizontal coordinate is X = (u − cx)Z/fx. Here u is a pixel coordinate and Z is depth. This equation describes camera geometry; it does not by itself identify the carton’s center, speed or future contact point. [4]

A second observation adds a time difference. After expressing both positions in a common frame, displacement divided by elapsed time gives an average velocity. Otherwise, the camera turning left can look like the box moving right. A longer sequence provides evidence through noisy samples or short occlusions, although stale observations still need a prediction of what happened since exposure. [4]

A learned intermediate vector should not be read as a list of named physical measurements. For interpretation, ask what it helps the controller do and which errors change its actions. A precise-looking box outline alone tells us little about interception timing. The next command must put the body where contact will happen, while accounting for its current posture.

Follow pixels through the robot perception pipeline

Reflex combines a visual motion estimate with robot state to command the research G1 body.
Original explanatory diagram of the documented input and control roles. It is not an experiment recording. Open the diagram for a larger view. [3]

Three training jobs share a control interface

The authors separate learning into three stages. First, reinforcement learning teaches catching from simulator state. Second, a compact representation links delayed, incomplete object observations to the controller. Third, visual histories learn to reproduce that representation while the body controller stays fixed. This separates visual estimation from discovering the catching motion. [1]

Stages 1 and 2 use human-motion references and an adversarial motion prior. Stage 2 initializes through teacher-action matching, aggregates visited states, then refines with proximal policy optimization (PPO). Stage 3 trains visual localization followed by temporal latent prediction. [5]

The engineering benefit is a narrower debugging question. A missed box could come from an incorrect motion estimate, from choosing a poor receiving posture, or from losing support after contact. Training the parts separately makes these interfaces available for controlled comparisons. It does not make an error in perception harmless; that error still changes the joint commands.

Catching continues after first contact

The project shows the G1 stepping toward off-center throws and receiving cartons across its forearms and torso. It also shows catching followed by carrying and handover. Those clips illustrate selected motions; the paper’s trial tables supply the success rates. [5]

Mechanically, moving the feet changes where ground forces act. Bending the legs moves the pelvis; rotating the torso changes arm reach and the position of the supported load. The shoulders and elbows must bring the forearms under the carton while the rest of the body manages the resulting contact forces. Each adjustment changes the state that the next control command must address. [6]

Consider a box reaching the right side of the chest. Extending only the right arm also shifts mass and leaves less surface under the box. A step can bring the body underneath, while coordinated arm motion provides support from both sides. This is an explanation of the mechanical requirements, not a claim that Reflex follows a fixed hand-written sequence. Its learned joint commands must handle these coupled effects. [6]

A contact frame is therefore weak evidence of a completed catch. The box may touch both arms and still slide out. For useful evaluation, keep interception, retained support and the robot’s final posture as separate outcomes.

Read each percentage with its trial setup

Reflex paper, Sections V-A and V-B, Tables II and VII
EvaluationProtocolResult
RGB-D simulation6 controllers × 640 throws3,290/3,840 = 85.7%
Physical G1 at 2.5 m20 throws13/20 = 65%
Physical G1 at 5 m20 throws8/20 = 40%
Physical G1 at 7 m20 throws3/20 = 15%

[3]

Simulation used 34 × 26 × 22 cm boxes from 3–3.5 m; physical distance trials used 31 × 26 × 21 cm boxes. Simulated catches require chest support and low speed for 0.6 s. [3]

With only 20 physical throws, one changed outcome moves the rate by five percentage points. The table supports a distance-dependent loss under this experiment. It does not establish an exact reliability figure for repeated warehouse work. Hand-thrown trajectories, box properties and reset conditions would need to match before comparing another controller’s percentage.

The project’s interactive 3D player replays recorded successful simulation runs. Its visual encoder was trained on throws extending to 7 m. Clicking through successful replays does not provide a fresh trial or a denominator for estimating reliability. [5]

Removing history changes the result by points

Table V, means across five policies per variant
Simulated throwFull historySingle frameDecrease
Centered85.2%49.6%35.6 points
0.6 m lateral offset66.6%26.8%39.8 points

[3]

The relative decreases are different calculations. For centered throws, 35.6 ÷ 85.2 is about 41.8%. For offset throws, 39.8 ÷ 66.6 is about 59.8%. Reporting a 35.6% relative reduction would misstate the first comparison. The larger offset loss also fits the geometry of the task. A sideways correction requires the robot to decide where to place support before the box reaches it.

These ablations test specified alternatives trained for this task. They support using motion history here, without proving that every recurrent network will outperform every single-image controller.

Where the evidence stops

Long-range visual grounding, size-blind control and omitted motion blur or exposure effects remain limitations. Hardware timing covers an older detector pipeline, excluding actuation; current visual latency is unresolved. [3]

A useful follow-up would report continuous exposure-to-actuation timing for the deployed visual model, including missed deadlines. Each failure should also identify whether the robot missed the flight path, contacted then dropped the carton, or retained it while taking an unstable posture. These measurements would separate a faster perception model from a better support strategy.

Different objects require separate trials. A deformable bag, a rotating tool and a rigid carton present different visible surfaces and contact behavior. The study does not provide evidence that the same controller can catch all three. Its strongest use for readers is showing how temporal perception and body control can be tested as connected components.

Read the existing G1 product profile

Compare catching with a planned grasp

Read the whole-body balance explanation

Research footage is supplied by the authors. Open the authors’ videos and project

Sources and verification

  1. Zhongzheng Ren publication record for Reflex ↗Zhongzheng Ren · Read 10 October 2026

    Author-maintained publication entry and abstract, with original paper and project links.

  2. Reflex code release status ↗Reflex research team · Read 10 October 2026

    Repository says training code, evaluation code, checkpoints and deployment stack are being prepared for release.

  3. Reflex scientific paper ↗Taoyang Jia and colleagues · Read 10 October 2026

    All 12 pages read. Figures 1–5 and Tables I–VII checked. Canonical PDF and the older project PDF are byte-identical.

  4. Camera calibration and coordinate geometry ↗OpenCV · Read 10 October 2026

    Pinhole projection and camera/world coordinate equations. Used for the independent geometric explanation, not for Reflex results.

  5. Reflex author project and videos ↗Reflex research team · Read 10 October 2026

    Canonical project page read in full. Videos remain external links because republication permission was not established.

  6. Multi-body dynamics and contact ↗Russ Tedrake, MIT · Read 10 October 2026

    Mechanical background for the independent discussion of coordinated joints and external contact.

  7. Unitree G1 and G1 EDU specification columns ↗Unitree Robotics · Read 10 October 2026

    Separate base G1 and EDU columns, joint counts and secondary-development support.

Article history

Added a technical reading of Reflex with separate simulation and physical protocols, visual-history arithmetic, source links and open evaluation questions.

Report a correction