ROBOTICS FIELD NOTESENGLISH EDITION / 8 October 2026
Articles

How Humanoid Robot Teleoperation Works

How VR tracking, gloves, exoskeletons and motion capture drive robot joints, with command rates, network delays, safety limits and learning data.

A person supplies the motion goal

Teleoperation lets a person control a robot through a remote interface. Tracking software measures a hand, headset or body movement. A mapping program converts it into a robot target. Local controllers move the joints, while cameras and other sensors return feedback. The human may choose every reach even while software handles joint limits and balance. Automatic joint control alone does not make the task autonomous.

Six ways to send an instruction

InterfaceWhat it measures or sendsDocumented example
VR headsetHead and wrist poses with a view from robot camerasOpen-TeleVision uses Apple Vision Pro and stereo video
Camera hand trackingFinger keypoints and wrist poseAnyTeleop uses RGB or RGB-D cameras
Motion captureTracked body posesTWIST uses an OptiTrack system for its G1 experiments
Data glovesPalm and finger poses, sometimes with touch feedbackShadow offers haptic and non-haptic glove options
JoysticksDirection, rotation or mode commandsUnitree XR documentation assigns separate movement and turning controls
Exoskeleton controllerAngles of an external linkage worn or moved by the operatorACE combines joint encoders with hand-facing cameras

[1] [2] [3] [4] [5] [6]

These interfaces can be combined. In ACE, the exoskeleton’s measured angles and link geometry give the wrist pose. A camera facing each hand estimates finger positions. The robot then solves for its own arm configuration. The operator’s linkage and the robot arm do not need matching joints. [6]

Fit a human movement to robot hardware

Retargeting maps a human pose to a pose the robot can achieve. Arm lengths, finger sizes and joint ranges differ. Copying human angles directly can put the robot wrist in the wrong place. Inverse kinematics finds robot joint angles for a requested hand position and orientation. Finger retargeting can minimise errors between corresponding fingertip vectors while respecting joint limits and limiting abrupt changes. [2]

The command path is human movement, tracking, retargeting, robot controller, joint motion and sensor feedback. Synchronisation requires timestamps. A left-hand pose and a right-hand pose captured at different moments can distort a two-hand action, such as holding a lid while turning its container. Record both observation time and command time when assessing the delay.

Human tracking and retargeting send goals to local robot control while sensor feedback returns to the operator.
Original teleoperation sketch. The network boundary shown is illustrative. Retargeting can run on either side, and joint control frequency is separate from the end-to-end delay. Open the diagram for a larger view.

Command rate is different from latency

A command rate counts updates per second. Latency is the time from an event to the response being measured. Camera exposure, pose estimation, network transport, buffering, motion calculation and video display all contribute. A fast stream can still contain old information. Variable delay also makes corrections arrive unevenly.

Measure command age and video age separately. Bandwidth limits may require a lower video frame rate or resolution. Packet loss can interrupt a stream, and buffering can leave the displayed image behind the robot’s current position. A return video alone therefore cannot confirm that a new stop command has arrived.

AnyTeleop receives end-effector targets at 25 Hz and generates joint trajectories at 120 Hz. Those figures describe two software stages in the research system. Its paper also reports tracking loss during fast hand movement and unreliable poses during self-occlusion. A command frequency does not measure task success or guarantee a particular internet connection delay. [2]

Open-TeleVision reports a 60 Hz loop with 480 × 640 images for each eye. The operator moves a head-mounted stereo camera and controls the arms and hands. Its H1 and GR-1 experiments use upper-body degrees of freedom; the paper does not evaluate walking during those tasks. [1]

Walking needs a controller under the operator

TWIST shows a different arrangement on a Unitree G1 with 29 degrees of freedom. OptiTrack captures human motion at 120 Hz. Retargeting and policy commands run at 50 Hz, while the robot’s proportional-derivative joint controller runs at 1,000 Hz. The authors roughly measured about 0.9 seconds of delay from video in this version of the system. That estimate covers observed teleoperation response, not network transport alone. [3]

TWIST trains a tracking controller in simulation using reinforcement learning and behaviour cloning. Its training teacher can inspect future motion frames; the deployed controller uses current reference motion and robot-state measurements. The paper shows a human directing G1 while it crouches to lift a box and moves objects with its feet. These examples are teleoperated. The Booster T1 results in the same manuscript are simulation tests. [3]

The separate article on humanoid walking control explains the balance constraints beneath a movement request.

Return contact to the operator

Bilateral teleoperation adds a mechanical feedback channel toward the operator, such as a resisting force. Tactile feedback can instead press or vibrate against the skin. A glove that measures finger position does not necessarily provide either sensation. The feedback hardware and the contact sensor need to be identified separately.

Shadow’s September 2025 specification describes fixed UR10e arms with Dexterous Hands. Its HaptX G1 option uses pneumatic tactile actuators in the gloves; the non-haptic option tracks movements without that touch output. These are robot-arm systems, not walking humanoids. Returning touch can help an operator notice contact, but it does not establish a safe force limit for an unknown object. [4]

Handle a lost stream before it becomes a motion

Teleoperation needs defined behaviour for missing tracking, stale commands and connection loss. A system can reject an unreachable target, limit motion speed or request a controlled stop. The stop strategy must suit the hardware. Removing torque from a standing humanoid can allow it to fall. A local emergency stop and a software pause address different failure paths.

Unitree’s XR repository documents a software damping command in motion mode, plus separate exit controls. The stated control bindings depend on the configured mode. Shadow lists physical emergency-stop buttons and a pedal, and monitors tendon forces. Neither description proves that every modified setup or network arrangement is safe. The integration needs its own stop and recovery checks. [5] [4]

Turn a session into training data

During data collection, save what the robot observed alongside the action sent to it. Useful records include camera images, joint positions, requested actions, timestamps and task outcomes. Mark interventions and interrupted attempts. Pairing an image with an action from the wrong moment teaches a different relationship from the one the operator intended.

ALOHA provides a fixed two-arm example. A person moves leader arms while the follower arms perform the task. Action Chunking with Transformers, or ACT, learns to predict sequences of actions from recorded examples. Training adjusts model parameters. During later execution, inference computes actions from new observations. Replaying the operator’s recording is a separate operation. [7]

Open-TeleVision also trained ACT-based policies. Its H1 can-sorting evaluation covered five episodes with ten cans each. The reported picking success was 92% and placement success 88%, assessed as separate stages. These are autonomous-policy results after training, not the operator’s teleoperation score. They do not establish uninterrupted success for the entire sorting sequence. [1]

Name who chooses the next action

  • Direct teleoperation means the person supplies the continuing motion requests.
  • Supervised control means a person assigns or approves goals and handles exceptions.
  • Partial autonomy means software completes a defined part of the task, such as an approach or grasp, within a wider human-controlled workflow.
  • Autonomous execution means the robot selects the task actions from its observations without live human motion commands during the reported run.

When assessing a video or a reported trial, use the questions in how to identify a robot’s operating mode. The source should disclose interventions, resets and the control interface used during that run.

For timestamp alignment and command-response comparison, use the physical-log calibration guide.

Sources and verification

  1. Open-TeleVision teleoperation with immersive active visual feedback ↗Xuxin Cheng, Jialong Li, Shiqi Yang, Ge Yang and Xiaolong Wang · Read 8 October 2026

    July 2024 manuscript. H1 and GR-1 upper-body setups, 60 Hz loop and separate autonomous-policy evaluation.

  2. AnyTeleop a general vision-based dexterous robot arm-hand teleoperation system ↗Yuzhe Qin and colleagues, RSS 2023 · Read 8 October 2026

    Primary paper. Separate 25 Hz pose input and 120 Hz motion generation. Lists hand tracking and occlusion failures.

  3. TWIST teleoperated whole-body imitation system ↗Yanjie Ze and colleagues · Read 8 October 2026

    May 2025 manuscript associated with CoRL 2025. Unitree G1 real-world teleoperation and Booster T1 simulation are distinct.

  4. Shadow Teleoperation System technical specification ↗Shadow Robot Company · Read 8 October 2026

    September 2025 specification. Fixed UR10e arms, haptic and non-haptic glove options, emergency stop hardware and tendon-force monitoring.

  5. Unitree XR teleoperation documentation ↗Unitree Robotics · Read 8 October 2026

    Repository checked on 8 October 2026. Model and mode-specific controls, data recording and a software damping command.

  6. ACE a cross-platform visual-exoskeletons system for dexterous teleoperation ↗Shiqi Yang and colleagues, UC San Diego · Read 8 October 2026

    August 2024 manuscript associated with CoRL 2024. Exoskeleton encoders measure wrist pose and hand-facing cameras measure fingers.

  7. Learning fine-grained bimanual manipulation with low-cost hardware ↗Tony Z. Zhao, Vikash Kumar, Sergey Levine and Chelsea Finn, RSS 2023 · Read 8 October 2026

    ALOHA is a fixed two-arm manipulation platform. ACT learns from recorded human control and predicts action sequences.

Article history

Technical explanation added with primary sources, hardware distinctions and experimental limits.

Added a contextual link to the simulation series for the next engineering step.

Report a correction