A pickup needs geometry and contact
A humanoid picks up an object by estimating where it is, choosing a reachable hand pose, moving the arm and closing the fingers. The grasp must produce enough contact force to support the object without damaging it. Camera images, joint measurements and, on some hands, touch sensors tell the controller whether the motion is working. These stages can use separate programs or a learned policy. A walking humanoid must also manage the change in body load.
Put the object in the robot’s coordinates
An RGB camera records colour and texture. Object detection identifies an image region that may contain a bottle; it does not by itself provide the bottle’s distance or orientation. A depth camera supplies range estimates. Pose estimation combines suitable measurements with object geometry to estimate position and rotation. Some planners predict a grasp directly from a depth image without identifying the object or estimating its complete pose. [1] [2]
The camera and arm use different coordinate frames. Hand-eye calibration estimates the fixed camera-to-hand relationship for a wrist-mounted camera. A separate arrangement estimates the relationship between a fixed camera and the robot base. OpenCV documents both. The robot combines that calibration with its measured joint angles to express the target in the arm’s frame. A loose camera mount can make a visually correct target produce a misplaced hand. [1]
Choose contacts and a path
A grasp pose specifies the hand’s position, orientation and finger opening. Candidate contacts must leave room for the fingers to close. Inverse kinematics calculates arm joint angles that can place the hand at the selected pose. The planner checks joint limits and collisions, then supplies intermediate targets for the approach. Joint controllers compare those targets with measured joint positions and adjust motor effort. A valid endpoint alone says little about whether the elbow can get there. [3] [4] [5]
For Open-TeleVision’s H1 demonstrations, researchers used the palm and all five fingers to hold cans for sorting. For insertion into slots, they used a pinch between thumb and index finger to leave room for adjustments. The same upper-body robot therefore used different grasps for different next steps. [6]
What closes the fingers
An electric motor creates torque. A transmission carries that torque to a finger joint. Tendon-driven designs pull cables around joints; some couple several joints to one actuator so the finger bends around an object. Direct-drive joints connect the motor to the joint without a reduction gearbox. Geared designs trade motor speed for joint torque, with friction and backlash that affect force estimates. These are different hardware choices, not features present in every hand. [4]
A parallel gripper moves two opposing jaws. An articulated hand has more contact choices, plus more joints to coordinate. Position control requests an angle or opening. Torque control requests rotational effort; tactile or force measurements can support corrections after contact. Motor current alone is an imperfect contact-force estimate when transmission friction is uncertain. [4]
How much force holds the object
Friction resists sliding at the finger surface. In a basic Coulomb contact model, available tangential force is bounded by friction coefficient multiplied by normal force. For an ideal vertical object pinched symmetrically between two fingers, 2μN must at least balance mg. Here N is the inward force per finger, μ the friction coefficient, m the mass and g gravitational acceleration. This static example omits acceleration, rotation and deformation. Raising the grip force may prevent slip while crushing a thin container. [3]
Read the contact before the object drops
Tactile sensors measure contact through signals such as pressure or deformation. Optical tactile sensors place a camera behind a soft surface. GelSlim tracks the contact imprint and markers in that surface. Different motion near the edge of the contact patch can reveal incipient slip, when only part of the contact is sliding. A controller can then adjust the grip or pause the movement. [7]
A 2018 MIT experiment used an ABB-1600 arm with a WSG-50 gripper. Researchers manually pushed, pulled and rotated ten held objects, including non-slip cases, over 240 tests. The detector achieved 86.25% accuracy. Flat, smooth contacts produced weak signals and missed slip events. That number measures detection accuracy, not successful pickups or humanoid task completion. [7]
The separate guide to robotic hand sensing compares touch measurements and the hardware that produces them.
Learn the motion and the corrections
Imitation learning fits a policy to recorded observations and actions. Teleoperation can supply examples of approaching, closing, lifting and recovering from a poor grip. Reinforcement learning instead adjusts a policy using rewards accumulated during interaction. A reward for lifting alone can miss damage or an unstable hold; those outcomes need their own measurements or constraints. [8] [9]
A vision-language-action model uses images and a language instruction to predict robot actions. RT-2 combined robot trajectories with vision-language training. Its authors state that web pretraining did not teach new physical motions by itself. Recognising a fragile glass in an instruction does not establish that the hand has learned a safe contact strategy. [10]
For the path from a person’s movement to a training record, read how humanoid teleoperation works.
Read the denominator
Dex-Net 2.0 gives a useful bounded example. On an ABB YuMi with silicone gripper tips, its GQ-CNN planner achieved 80% success across 50 trials on ten novel household objects. That is 40 successful attempts. Objects were isolated on a table, with a person resetting their pose between trials. Success required lifting, transporting and retaining the object after shaking. These were fixed robot-arm tests, without a humanoid balancing or walking. [2]
The same experiment reported 100% precision among 29 grasps classified as robust. That subset statistic does not mean all 50 attempts succeeded. Keep total attempts, selected subsets and complete task cycles separate when comparing grasping claims. [2]
Where a pickup breaks down
Transparent objects can leave holes or background measurements in a depth image. ClearGrasp addresses this by estimating transparent surfaces, their orientation and occlusion boundaries before correcting the depth map. This changes the geometry supplied to a grasp planner. It does not establish the object’s weight, friction or resistance to crushing. [11]
- Deformable objects change shape as the fingers close, so a pose estimated before contact can become stale.
- Slippery surfaces reduce the tangential load that a given squeeze can support.
- Small parts leave little contact area and little clearance between the finger and the table.
- Fragile objects impose a force limit that may conflict with the grip needed for a fast lift.
A useful trial log records the object, starting pose, surface condition, hand configuration, failed stage and any human intervention. It should also state whether success ends at the first lift or after placement. Those details show which part of the pickup chain still needs work.
Sources and verification
- Camera calibration and 3D reconstruction ↗OpenCV · Read 8 October 2026
Official hand-eye calibration documentation. Camera-to-gripper and camera-to-base mappings require distinct configurations.
- Dex-Net 2.0 deep learning to plan robust grasps ↗Jeffrey Mahler and colleagues, UC Berkeley · Read 8 October 2026
RSS 2017 paper. Table IV reports 80% across 50 trials on ten novel tabletop objects, with human resets between attempts.
- Robotic Manipulation chapter on bin picking ↗Russ Tedrake, MIT · Read 8 October 2026
Course notes on collision geometry, contact forces, friction and grasp selection. Equations are models with stated assumptions.
- Robotic Manipulation chapter on robot hardware ↗Russ Tedrake, MIT · Read 8 October 2026
Position and torque control, transmissions, parallel grippers and underactuated fingers.
- Basic Pick and Place ↗Russ Tedrake, MIT · Read 8 October 2026
Course notes on frames, trajectories and differential inverse kinematics.
- Open-TeleVision teleoperation with immersive active visual feedback ↗Xuxin Cheng, Jialong Li, Shiqi Yang, Ge Yang and Xiaolong Wang · Read 8 October 2026
July 2024 manuscript. H1 and GR-1 upper-body setups, 60 Hz loop and separate autonomous-policy evaluation.
- Maintaining Grasps within Slipping Bound by Monitoring Incipient Slip ↗Siyuan Dong, Daolin Ma, Elliott Donlon and Alberto Rodriguez, MIT · Read 8 October 2026
2018 manuscript. ABB-1600 arm and WSG-50 gripper. Slip detection accuracy is separate from pickup success.
- Key concepts in reinforcement learning ↗OpenAI · Read 8 October 2026
Definitions of policies, observations, actions and cumulative rewards. No robot-specific performance claim.
- Learning fine-grained bimanual manipulation with low-cost hardware ↗Tony Z. Zhao, Vikash Kumar, Sergey Levine and Chelsea Finn, RSS 2023 · Read 8 October 2026
ALOHA is a fixed two-arm manipulation platform. ACT learns from recorded human control and predicts action sequences.
- RT-2 vision-language-action models transfer web knowledge to robotic control ↗Anthony Brohan and colleagues, Google DeepMind · Read 8 October 2026
2023 research manuscript. Language and image inputs condition robot actions; the authors describe limits on physical skills.
- ClearGrasp research project ↗Shreeyak S. Sajjan and colleagues · Read 8 October 2026
ICRA 2020 project. RGB-D depth correction for transparent objects. No claim of universal transparent-object handling.
Article history
Explained inverse kinematics and joint feedback, added direct depth-based grasp planning and compared two H1 hand grips from Open-TeleVision.
Report a correction