Skip to main content

View as Markdown

Page content converted to Markdown. Use the original page link at the end to explore interactive graphics.

1. What is embodied AI

1.1 Definition​

Embodied AI integrates artificial intelligence into physical systems, such as robots and intelligent vehicles. Through a physical body, AI uses a closed loop of perception, decision making, and action to interact dynamically with the real world.

The definition has three key ideas:

  • Body: the agent has sensors and actuators, so it can both observe and change its environment. The body is usually a real robot; in research and training, it can also be a virtual robot in simulation.
  • Interaction: the agent's actions change the environment, and those changes determine what it observes next.
  • Closed loop: perception, decision making, and action repeat continuously. The agent does not simply look once and produce an answer.

1.2 The perception–decision–action loop​

For a robot playing soccer, the goal is to score while keeping its body balanced:

Embodied AI · Information and action

From observing the world to changing it

Select a stage to explore its inputs, processing, and outputs.

Continuous loop

Inside the robot

External world

Perception and understanding

What is the state of the world and the robot?

Input
Camera images, joint angles, and body pose.
Processing
Detect the ball, goal, and goalkeeper; estimate their positions relative to the robot.
Output
An estimate of the scene and the robot’s state.

On the soccer fieldSee the ball ahead and the goalkeeper covering the right side of the goal.

A view of information flow: in a real robot, these stages run continuously at different rates.

The loop does not run just once. The robot repeats it throughout the kick, and each stage runs at its own pace: the vision module that detects the ball and goal may update dozens of times per second, while joint controllers may update hundreds to thousands of times per second.

In the 3D demo below, you can rotate the view and change the playback speed to follow the loop step by step:

Robot soccer · A 3D feedback loop

Teach a robot to score a goal

Shooting robotGoalkeeper

Preparing the 3D field…

New observations return to perception

PerceptionThe camera captures the field and identifies the ball, goal, and goalkeeper.
Play the demo or select a stage to explore it step by step. Moving the goalkeeper restarts the demo. Watch how the shooting direction changes. Motion is simplified to illustrate the feedback loop.

"Perception, decision, and action" is a functional view of the system. A real system may use several collaborating modules or a single model with multiple roles.

1.3 A body changes the problem​

Many AI systems work with data that has already been collected: an image goes in and a label comes out; a passage of text goes in and an answer comes out. An embodied agent has to act in its environment, which brings several fundamental differences:

AspectAI that processes static dataEmbodied AI
Where data comes fromExisting datasets that the model receives passivelyThe agent's own actions determine what it observes next
What it outputsLabels, text, or imagesActions that affect the physical world
Cost of mistakesUsually, you can just generate againThe robot may fall, break objects, or collide, often irreversibly
TimingCan compute offline; being slower is acceptableMust act within each control cycle, or the robot loses balance or misses its chance
Scale of dataHuge amounts of text and images on the internetRobot interaction data must be collected specifically; it is expensive and scarce

These differences explain why large models improve quickly at conversation and writing, while getting robots to fold laundry or do housework reliably remains hard. Chapter 5 discusses these problems one by one.

Moravec's paradox​

In 1988, roboticist Hans Moravec observed that it is comparatively easy to make computers perform at an adult level on intelligence tests or at checkers, but difficult or impossible to give them the perception and mobility of a one-year-old. This observation became known as Moravec's paradox.

Moravec's explanation was that perception and movement are the product of an extremely long evolution. They are so highly optimized that people perform them effortlessly and are unaware of how difficult they are. Abstract reasoning appeared much later and is easier to write down as explicit rules. For embodied AI, this means that actions that look simple are often the hardest, such as picking up a cup the robot has never seen or keeping balance on a slippery floor.

1.4 The body shapes what can be learned​

The same algorithm can do very different things on different bodies:

  • Sensors determine observations: with only an overhead camera, a robot arm may not see details near its gripper; adding a wrist camera makes close-up manipulation easier.
  • Actuators determine actions: a parallel gripper suits regularly shaped objects, while in-hand manipulation such as spinning a pen or twisting a bottle cap needs a dexterous hand.
  • Form determines where the robot can go: a wheeled base is efficient on flat ground but cannot climb stairs; a legged robot can step over obstacles but is harder to control.

So before discussing an embodied AI method, state which body it runs on. Chapter 4 introduces common robot platforms.

1.5 Embodied AI and robotics​

Robots are the most important platform for embodied AI. The two fields overlap heavily but emphasize different things:

  • Robotics studies the design, modeling, perception, planning, and control of robots, and provides foundations such as kinematics, dynamics, state estimation, and control theory.
  • Embodied AI emphasizes the interaction of intelligence, body, and environment. It focuses on how agents learn to complete tasks through interaction and generalize to new objects, environments, and instructions.

Today, embodied AI research usually combines the modeling and control methods of robotics with the data-driven methods of machine learning. The foundations on this site cover both.

Besides robots, autonomous vehicles, drones, and virtual agents that learn navigation and manipulation in simulation are also often considered part of embodied AI. This site focuses on robots.

Summary​

  • Embodied AI studies how agents use a body to interact with the environment and complete tasks through a closed loop of perception, decision making, and action.
  • Compared with AI that processes static data, embodied AI must gather information actively, bear the physical consequences of its actions, meet real-time requirements, and cope with scarce data.
  • The body determines observations, actions, and reachable environments. When discussing a method, first state which body it uses.

Further reading​

References​

  1. NVIDIA: What is embodied AI?
  2. Hans Moravec. Mind Children: The Future of Robot and Human Intelligence. Harvard University Press, 1988.