Skip to main content

View as Markdown

Page content converted to Markdown. Use the original page link at the end to explore interactive graphics.

3. Robot skills

To study embodied AI, start by defining the task. Modeling the task well and choosing appropriate goals and evaluation metrics are often more important than model selection and optimization. This chapter first explains how to define a task, then introduces four common robot skills.

3.1 Define the task first​

Before choosing a method, clarify the desired skill, environment, and constraints. Which sensors will you use? What are the input observations and output actions? What counts as completion, and how will you measure performance through success rate, completion time, or other metrics?

Also examine whether the task setup is appropriate and whether changing it could help. For example, if an overhead camera cannot observe enough of the environment, a wrist camera may provide a useful close-up view. Changing sensors changes the observation conditions: document this explicitly and compare methods under consistent settings.

3.2 Turn the task into a learnable problem​

Before handing a task to an algorithm, write it down as explicit inputs, outputs, and goals. A common formulation: at each time step , the robot receives an observation , and a policy outputs an action based on the observation and the task goal (for example, a language instruction):

The action affects the environment, the environment changes, the robot receives a new observation, and the loop continues. Reinforcement learning also defines a reward that measures how well each step went. For formal definitions of rewards, returns, and value functions, see Markov decision processes.

The table below illustrates these elements with two tasks:

ElementMeaningRobot arm picking up a cupQuadruped walking forward
Observation Information available to the robotCamera images, joint angles, gripper openingJoint angles and velocities, orientation and angular velocity from the IMU, velocity command
StateComplete information about the world, often not directly observableThe cup's true pose, mass, and frictionGround friction, foot contacts, the body's true velocity
Action The policy's outputEnd-effector pose increments or joint targets, plus gripper open/closeTarget angles for 12 joints
Goal What to accomplishPlace the cup in the target areaWalk forward at a given velocity
Success conditionWhat counts as doneThe cup is placed correctly and upright, and the arm avoids collisionsTrack the target velocity, stay balanced, and slip less
Control frequencyHow often an action is producedFrom a few hertz to tens of hertzFor example, a 50 Hz policy, with faster low-level joint control

Three points are worth noting:

  • Observations often reveal only part of the state. A camera cannot see how heavy a cup is; the robot can only infer it from forces and motion after picking the cup up. This is called partial observability. Policies often compensate by using past observations or adding sensors.
  • The action definition affects how hard learning is. Joint targets or end-effector poses, absolute values or increments, position control or torque control: each has its trade-offs.
  • Compare methods under consistent settings. When methods use different observations, actions, control frequencies, or success conditions, their results cannot be compared directly.

3.3 Four common skills​

A useful starting point is four common skill categories: grasping, manipulation, locomotion, and navigation. Grasping and manipulation change object states; locomotion and navigation move the robot itself. Each skill comes with a 3D demo you can play step by step.

Grasping​

Key question: how do you hold an object securely?

Grasping is the process and result of establishing contacts between an end effector and an object so that the object's motion is constrained, it can move with the end effector, and it does not slip under expected disturbances.

In everyday terms, a robot uses a gripper, dexterous hand, or suction cup to hold an object securely before moving or manipulating it. To put a small block into a box, it first grips and lifts the block. Establishing and maintaining that hold is grasping; transporting, placing, and releasing it complete the pick-and-place task.

3D skill demonstration

Hold securely, then lift

Preparing the 3D scene…An arm aligns its parallel gripper with an orange block, closes the fingers, and lifts the block off the table.

Observe the objectFind the block’s position and orientation, then choose contact regions on opposite sides.

Watch the gripper–object relationship change from no contact to a stable grasp.Drag the progress slider or select a stage to explore step by step. Motion is simplified to illustrate the skill.

Typical grasp detection takes point clouds or RGB-D images as input and predicts a grasp pose for a parallel gripper: 3D position, 3D orientation (6-DoF), and gripper width. Execution also requires reachability and collision checks.

Geometric and mechanical approaches analyze conditions such as form closure and force closure. Data-driven methods learn grasp quality or generate grasp poses directly. Representative projects include Dex-Net, the dataset and benchmark GraspNet-1Billion, and AnyGrasp. Analytical models can also generate training data.

A grasp can be planned before execution or continuously corrected through vision, force, or touch; AnyGrasp, for example, supports grasp pose tracking. Grasping regular rigid objects in controlled workcells is relatively mature. Dexterous hands, transparent or deformable objects, heavy occlusion, and clutter remain reliability challenges.

Manipulation​

Key question: how can contact change an object's state?

Manipulation is the purposeful application of forces or motion to change an object's position, orientation, shape, or other task-relevant state.

In everyday terms, a robot pushes, pulls, carries, or rotates something into the desired state. For example, it grasps a drawer handle and pulls along the rails until the drawer reaches a target opening.

3D skill demonstration

Grasp the handle and open the drawer

Preparing the 3D scene…The arm grasps the drawer handle, maintains contact, and pulls the drawer along its rails.

Approach the handleObserve the handle and rail direction, then move the gripper toward the handle.

Grasping is only the start. Manipulation continues to change the drawer’s state.Drag the progress slider or select a stage to explore step by step. Motion is simplified to illustrate the skill.

Manipulation is broader than grasping, which is one of its subproblems. A stable grasp usually maintains the end-effector–object contact relationship so the object moves with the end effector. Manipulation may change contacts repeatedly, such as switching fingers, rolling, or regrasping during in-hand rotation. Fixed contact points are not an absolute boundary between grasping and manipulation.

Manipulation includes grasping, transporting, pushing, nudging, flipping, and in-hand manipulation, as well as folding clothes, arranging ropes, opening doors and drawers, using tools, bimanual coordination, and long-horizon tasks. Pushing an object or kicking a ball is non-prehensile manipulation: neither requires first grasping the object.

Complex manipulation often requires continuously choosing actions from images, proprioception, and optional language goals. Imitation learning (IL) and vision-language-action models (VLA) are important research directions: ACT predicts action chunks; Diffusion Policy combines action generation with receding-horizon execution; π0 and OpenVLA explore multitask vision-language-action policies.

Challenges include rich and changing contacts, partially observed states, multiple valid actions for one goal, and errors accumulating over long sequences. Closed-loop feedback matters for complex manipulation and can also be used for grasping. Distinguish reproducing a fixed task from generalizing to new objects and environments.

Locomotion​

Key question: how can the body move stably?

Locomotion is the process by which a robot moves itself through interaction between mechanisms such as legs or wheels and the environment, while coordinating body posture and motion stability.

In everyday terms, it is about getting the robot's body moving and keeping it stable. A quadruped climbing a step must coordinate leg lifts, foot placement, and weight shifts to avoid tripping or losing balance. To choose its own route and reach a doorway, it must also coordinate with navigation.

3D skill demonstration

Stay balanced and climb the steps

Preparing the 3D scene…A quadruped alternates leg lifts, foot placement, and support to climb three steps.

Prepare supportAll four feet support the body as the robot prepares to move forward.

Watch how the body moves by coordinating foot placement, joints, and center of mass.Drag the progress slider or select a stage to explore step by step. Motion is simplified to illustrate the skill.

Legged robots must handle gaits, balance, steps, slopes, and fall recovery. Wheeled robots also perform locomotion, dealing with steering constraints, slipping, and uneven terrain. Whole-body control can further connect mobility with manipulation.

Observations often include joint positions and velocities, orientation and angular velocity from an inertial measurement unit (IMU), and optional depth or elevation maps. A policy may output joint targets for a low-level controller to track, or output torques directly. Distinguish the gait policy update rate from the joint servo rate.

Model predictive control (MPC) and sim-to-real reinforcement learning are common approaches. Learning to Walk in Minutes trains ANYmal through massively parallel simulation in Isaac Gym. Domain randomization helps bridge reality gaps; teacher–student training can transfer skills learned with privileged information to onboard observations. Work such as Robot Parkour further explores using vision to select and execute obstacle-crossing skills.

Quadrupeds have demonstrated strong locomotion in many tested settings. Extreme terrain, low-friction contacts, long-term reliability, and whole-body coordination for humanoids or changing payloads still require validation under specific conditions.

Key question: where to go, which route to take, and when to stop?

Navigation is the process of choosing and adjusting a robot's direction or path using its goal, environmental observations, and estimated state, while avoiding obstacles and reaching a target location or area.

In everyday terms, navigation answers: where am I, where should I go, and how do I get there? A robot might cross a room, avoid obstacles on the floor, and reach a designated area at the doorway.

3D skill demonstration

Avoid obstacles and reach the goal

Preparing the 3D scene…A wheeled robot follows a planned route around two obstacles and stops at the goal.

Locate the robot and goalIdentify the current position, goal, and obstacle locations.

Watch the goal and route: where to go, how to avoid obstacles, and when to stop.Drag the progress slider or select a stage to explore step by step. Motion is simplified to illustrate the skill.

Navigation chooses the destination and route; locomotion determines how the body moves along that route. In a typical layered system, navigation provides paths, waypoints, or desired velocities. Motion control coordinates gaits, footholds, joints, or wheel speeds to execute them. For a quadruped heading to a doorway, navigation may choose a route around boxes and over steps; motion control handles leg lifts, foot placement, and balance on those steps.

The two must continually cooperate. Planning must consider whether the robot can pass a narrow opening or climb a step. If it slips or is blocked, state feedback prompts a speed adjustment or replanning. This coordination can use separate modules, joint planning, or learned policies.

Navigation involves localization, mapping, path planning, and obstacle avoidance. Classical systems often combine localization / SLAM, global planning such as A*, and local planning or control such as DWA or MPC. Nav2 is a navigation framework covering these components. Mature engineering solutions exist for known maps and controlled environments.

Embodied navigation tasks can also be classified by goal: PointNav reaches specified coordinates; ObjectNav searches for an object category; ImageNav finds a place or object matching an image; VLN follows natural-language instructions. Inputs may include RGB-D, LiDAR, odometry, and goal descriptions. Outputs may be velocity commands or discrete actions such as move forward, turn, and stop.

Some simulation benchmarks simplify low-level movement. Real deployment must also handle localization errors, slipping, dynamic obstacles, and execution failures. Compare task definitions, sensors, success conditions, and path efficiency together; a high success rate on one PointNav benchmark does not generalize to all navigation tasks.

3.4 Compare the skills​

A frequency describes the update rate of a particular layer. The examples below come from specific systems. Check policy inference, action execution, simulation stepping, and joint servo rates separately.

SkillTypical inputsTypical outputsExample timescaleCommon methodsProgress and limitations
GraspingPoint clouds / RGB-D; optionally force and touchGrasp pose, gripper width, or contact parametersTriggered by the task or updated continuouslyGeometry and mechanics, supervised learningMature in controlled workcells; complex objects and contacts remain challenging
ManipulationImages and proprioception; optionally language, force, and touchEnd-effector or joint actions, action chunksACT example: actions executed at 50 HzPlanning and control, IL, RL, VLAVaries by task; general and long-horizon manipulation remain challenging
LocomotionJoint states and IMU; optionally terrain sensingJoint targets, torques, or foot targets50 Hz policy / 200 Hz simulation control (derived from the legged_gym base configuration)MPC, whole-body control, sim-to-real RLDemonstrated in many quadruped settings; complex terrain and whole-body coordination remain challenging
NavigationVision / LiDAR, pose estimates, goal descriptionsPaths, waypoints, velocities, or discrete actionsHigh-level updates as needed; Nav2 local control defaults to 20 HzLocalization / SLAM, planning and control, learned methodsMature with known maps; open environments and semantic goals remain challenging

3.5 Combine skills for a real task​

Real tasks often combine several skills:

  • Mobile manipulation: connects mobility with manipulation. The robot usually navigates to a work location, then coordinates its base and arm to complete the task. For example, "Go to the kitchen and bring back the cup."
  • Loco-manipulation: emphasizes the coupling of motion, balance, and contact-based manipulation, such as an armed quadruped carrying an object or a humanoid opening a door while balancing. ALMA is one example of legged whole-body manipulation.

A humanoid performing everyday tasks may use all four capabilities. What it can do depends on hardware, perception, control, and training, as well as its physical form.

Summary​

  • Studying embodied AI starts with defining the task: the skill, environment and constraints, observations and actions, and success conditions and evaluation metrics.
  • Grasping and manipulation change object states; locomotion and navigation move the robot itself. The four skills differ widely in inputs, outputs, and timescales.
  • Real tasks often combine several skills. Compare methods under consistent settings.

Further reading​

References​

  1. Modern Robotics: Grasping and manipulation
  2. Modern Robotics: Wheeled robots and mobile manipulation
  3. Habitat: ObjectNav and ImageNav tasks and evaluation
  4. VLN-CE: Vision-and-language navigation in continuous environments