ROBOTICS FIELD NOTESENGLISH EDITION / 8 October 2026
Articles

How Do Humanoid Robots Think and Make Decisions?

Follow a bottle-to-table task through perception, planning, learned policies, joint control, result checks and stops, using documented research architectures.

A decision selects the next action

A humanoid makes a computational decision when software selects a task, movement or stop condition from its current inputs. The selection may follow programmed rules, a planner or a learned policy. It does not establish human-like thought. Perception estimates the scene, planning chooses a route to the goal, and feedback control adjusts physical motion. [1]

Task planning handles the order of operations, such as approach, grasp, carry and release. Motion planning searches for movements that respect geometry and joint constraints. A policy maps observations, and sometimes an instruction or history, to actions. Low-level control compares requested motion with measured joint behaviour. These functions can sit in separate modules or share a learned model. [1]

Follow one bottle to a table

StageInformation or actionReason to pause or stop
1. Receive the instructionExtract the requested object and destination. A voice interface first converts speech to text.Several bottles or tables match the request. Ask for clarification.
2. Identify the bottleMatch the object description to a visible region.The target is missing or the match is uncertain. Obtain another view.
3. Locate itEstimate bottle pose and tabletop height in the robot’s working frame.Depth is absent or the estimate is stale. Do not start the reach.
4. Plan the approachChoose a reachable stance and a hand path around obstacles.No checked path exists. Reposition within the allowed workspace or stop.
5. Choose a graspSelect finger contacts that fit the bottle and the available hand.The bottle is outside the hand’s opening or permitted load. Reject the grasp.
6. ExecuteMove the arm, close the fingers, lift and carry while reading feedback.Unexpected contact or grip loss triggers the configured recovery or safe stop.
7. Verify placementCheck that the bottle is supported by the table before release and inspect it afterward.An issued release command alone does not confirm placement.
8. Correct or finishReobserve after a bounded recovery attempt and report the result.Repeated uncertainty ends the attempt and requests human assistance.

A motion planner can return a geometrically valid path without a motion schedule. Time parameterization adds velocity and acceleration constraints. MoveIt documents this separation. Moving a filled bottle also changes the carried load and the geometry that collision checks must consider. The example needs a controller and robot model configured for those conditions. [2]

The measurement errors in the third stage are explained in how humanoid robots see the world. Depth and object identity need separate checks.

Instruction, approach planning, joint execution and result verification for a bottle placement, with a correction branch.
An illustrative bottle task. This sequence explains separate responsibilities and does not report a completed experiment or imply every robot uses four separate programs. Open the diagram for a larger view.

Language scores meet physical skills

SayCan offers a documented task-selection example. Its language model scores the usefulness of available skills, while learned value functions estimate their chance of completion in the current state. The combined score selects a skill. The 2022 experiments used an Everyday Robots mobile manipulator with a seven-joint arm and a two-finger gripper. These results concern a wheeled platform. [3]

The tested system mixed learned and programmed skills. Picking used a learned policy and value function. Navigation used known object locations and a classical planner. Placement used motion planning and released the object after contact with a supporting surface. [3]

Across 101 instructions, PaLM-SayCan achieved 84% planning success and 74% execution success in its mock kitchen. In the real office kitchen, those figures were 81% and 60%. Human raters assessed plans and executions separately. The gap shows why a plausible sequence cannot establish that an object reached its destination. [3]

Where learned models enter

A vision-language model, or VLM, processes images and text. Its outputs can describe a scene or support a task choice. A vision-language-action model, or VLA, also produces robot action representations. Those outputs can be joint targets or hand movements, depending on the system. A readable answer and an executable command have different interfaces. [4]

NVIDIA’s March 2025 GR00T N1 preprint connects an Eagle-2 VLM to a diffusion-transformer action module. The latter uses visual and language features alongside robot state to generate action chunks. NVIDIA reports sampling 16 actions in 63.9 ms on an L40 GPU using bf16 precision. That measurement covers action generation on the stated GPU. The real-robot work used Fourier GR-1 for tabletop manipulation; the report lists short task horizons as a limitation. [4]

For the model terminology and its limits, see vision-language-action models explained. A VLA is one available architecture, and does not identify the software inside an unnamed robot video.

Learning and memory have separate jobs

Imitation learning trains a policy from recorded behaviour. In behaviour cloning, training adjusts the mapping from observations to recorded actions. A teleoperated reach can therefore become one training example. During later autonomous execution, the policy must handle its own errors and the observations they cause. States absent from the training data can lead to failed recovery. [5] [6]

Reinforcement learning adjusts a policy through rewards associated with outcomes. SayCan used learned skills and reinforcement-learning value functions for feasibility estimates. Those training choices are documented for SayCan. Other humanoid systems can use different methods. [3]

Robot memory can mean a map, a recent observation sequence or a record of attempted tasks. The storage must match the question the controller needs to answer. A previous bottle position is useful only with its age and reference frame. Recording that a grasp was requested must remain separate from evidence that it succeeded.

The distinction between training examples and reward-based learning is developed in how robots learn from actions and rewards.

Predicted futures can still be wrong

A world model predicts how a scene or internal state may change. Some systems use those predictions to plan, generate training data or derive actions. In its 12 January 2026 report, 1X describes 1XWM generating a future video for NEO, followed by an inverse-dynamics model that extracts an action trajectory between frames. [7]

1X reports 11 seconds for the video backbone to generate five seconds of footage, followed by one second for action extraction. It also describes depth errors that make NEO stop short of, or move beyond, a target despite a plausible generated video. Faster reactions and longer tasks with memory were listed as further work. The report describes the manufacturer’s own experiments. [7]

Control must keep checking

A controller needs fresh joint measurements while the arm moves. MoveIt Servo accepts joint velocity, hand velocity or hand pose commands and can check joint limits, collisions and singularities. A singularity is a configuration where some hand motions require excessive joint motion. Its collision checks can reduce commanded speed, and the documentation states that those checks are optional. Configuration therefore matters. [8]

The bottle example also requires a stop policy that remains available while a language model is busy. Delayed images, network loss and a new obstacle can invalidate the previous action. A stop instruction must reach the configured robot controller. Its physical response depends on the hardware and operating mode, including whether releasing a held object would create another hazard.

Measure the complete attempt. Record command-to-motion delay, successful placements, retries and human interventions under the same object and workspace conditions. That record distinguishes an accepted instruction from a task the robot completed.

Sources and verification

  1. Robotic Manipulation Perception, Planning and Control ↗Russ Tedrake, MIT · Read 8 October 2026

    Course introduction separates semantic and geometric perception, task planning, physical constraints and contact control.

  2. MoveIt motion planning concepts ↗MoveIt project · Read 8 October 2026

    Official documentation for joint constraints, collision checks and time parameterization of motion paths.

  3. Do As I Can, Not As I Say Grounding Language in Robotic Affordances ↗Michael Ahn and colleagues, Robotics at Google and Everyday Robots · Read 8 October 2026

    2022 research paper. Table 2 gives 84% planning and 74% execution in the mock kitchen, versus 81% and 60% in the real kitchen, across 101 instructions. A mobile manipulator was used.

  4. GR00T N1 An Open Foundation Model for Generalist Humanoid Robots ↗NVIDIA · Read 8 October 2026

    March 2025 technical report preprint. Sections 2 and 4.6 specify the VLM/action architecture, 63.9 ms action-chunk sampling on an L40 GPU, and short tabletop-task scope.

  5. Imitation Learning ↗Russ Tedrake, MIT · Read 8 October 2026

    Course notes on behavior cloning, observation histories and distribution shift. The article uses the supported definitions, not unfinished examples in the notes.

  6. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning ↗Stephane Ross, Geoffrey Gordon and Drew Bagnell, AISTATS 2011 · Read 8 October 2026

    Original paper explains how earlier actions change the observations encountered later in imitation learning.

  7. 1X World Model from Video to Action ↗1X Technologies AI Team · Read 8 October 2026

    Manufacturer research report dated 12 January 2026. Distinguishes generated video from executed motion and reports inference latency, depth errors and future memory work.

  8. MoveIt Realtime Servo ↗MoveIt project · Read 8 October 2026

    Official documentation describes joint and hand commands, joint limits, singularity checks and optional collision checking. Motion smoothing is also optional.

Article history

Specified SayCan’s learned and programmed skills, corrected the Servo source note and added the original DAgger reference.

Report a correction