Skip to main content

View as Markdown

Page content converted to Markdown. Use the original page link at the end to explore interactive graphics.

2. A brief history

Embodied AI is not a sudden new idea. Its core ideas go back to early research in artificial intelligence and robotics; its current momentum comes from deep reinforcement learning, large-scale simulation, and large models together. This chapter traces that history around one question: where do a robot's capabilities come from?

Four periods at a glance​

PeriodTimeCore approachRepresentative work
Programmed rules1960s–1990sPeople write the rules: first symbolic reasoning, then reactive rules in behavior-based robotics, which introduced the idea of embodimentShakey, A*, STRIPS, subsumption architecture
Model-based control1990s–2010sPhysical and probabilistic models for estimating state, planning paths, and controlling motionSLAM, Stanley, ZMP, MPC, ROS
Deep reinforcement learning2012–2021Learn control policies by trial and error in simulation or the real world; deep learning also improves perceptionEnd-to-end visuomotor policies, QT-Opt, the Rubik's cube hand, ANYmal, legged_gym
Foundation models2022–presentPretrained large models bring language understanding and general capabilities, working with low-level controlSayCan, RT-2, Open X-Embodiment, π0

Each period marks the main new source of capability at the time; it does not mean earlier methods became obsolete. Robots still rely heavily on planning and control today; legged locomotion is still mostly trained with deep reinforcement learning, and VLAs have started to use reinforcement learning to keep improving from deployment experience.

2.1 Programmed rules (1960s–1990s)​

In this period, robot capabilities came from rules written by people: first symbolic and logical reasoning, then rules that react directly to the environment, which introduced the idea of embodiment.

Symbolic reasoning: sense–plan–act​

From 1966 to 1972, SRI (then the Stanford Research Institute) built Shakey, widely regarded as the first mobile robot able to reason about its own actions. Shakey perceived its surroundings with a camera and a rangefinder, represented the environment as symbols, and used a planner to search for action sequences. The A* search algorithm and the STRIPS planner, both developed for Shakey, remain foundations of path planning and task planning today.

This pipeline, in which a robot first perceives, builds a world model, plans, and finally executes, is known as the sense–plan–act paradigm. Its problems: whenever the environment changes, the world model must be rebuilt; planning is slow, so the robot often stops to "think" for a long time; and a symbolic world model struggles to capture the full detail of the real world.

Behavior-based robotics and embodiment​

In the 1980s, Rodney Brooks at MIT proposed a different approach. His subsumption architecture does not build a complete world model. Instead, it organizes behavior into layers: lower layers handle reflexes such as obstacle avoidance, and higher layers suppress or invoke the lower ones when needed. The robot reacts directly to sensor signals, which makes it faster and more robust in the real world.

Brooks argued that intelligence can emerge directly from an agent's interaction with the real world: rather than maintaining a world model inside the robot, treat "the world as its own best model." Around the same time, researchers found that the body itself can take over part of the "computation": a passive dynamic walker, with no motors or controllers, can walk steadily down a gentle slope using only the structure of its legs and gravity.

From then on, embodiment became an important theme in artificial intelligence and cognitive science: intelligence lives not only in algorithms, but also depends on the body and the environment.

Whether symbolic or reactive, however, the rules still had to be designed by people, and robots struggled with situations their designers had not anticipated.

2.2 Model-based control (1990s–2010s)​

In this period, robot capabilities came from mathematical models of the world: probabilistic models describe sensor noise and uncertainty in the environment, and physical models describe how the robot moves. Here, "model" means a physical or probabilistic model, not the large models of the foundation model period.

Real sensors are noisy, and the environment is never fully known. From the 1990s, probabilistic methods became mainstream in robotics: Kalman filters and particle filters estimate the robot's state, and simultaneous localization and mapping (SLAM) lets a robot build a map of an unknown environment while locating itself in it. In 2005, Stanford's autonomous car Stanley completed a desert course of about 212 km to win the DARPA Grand Challenge, demonstrating the power of probabilistic perception and planning.

On the control side, model-based methods matured: the zero moment point (ZMP) method enabled humanoids such as Honda's ASIMO to walk, and model predictive control (MPC) and whole-body control allowed legged robots to run, jump, and resist disturbances. The Robot Operating System (ROS), whose development began in 2007, provided common communication mechanisms and tools for robot software and greatly lowered the barrier to building systems.

Systems of this period were mostly modular: perception, localization, planning, and control were developed separately with clear interfaces, and worked reliably in structured environments. The cost was that every module required extensive manual design and accurate modeling, and systems struggled to adapt to new objects, new environments, or contacts that are hard to model.

2.3 Deep reinforcement learning (2012–2021)​

Around 2012, deep learning achieved breakthroughs in image recognition, and robot perception improved dramatically. The bigger change came in control: with deep reinforcement learning, robots learn control policies themselves through repeated trial and error in simulation or the real world, instead of relying on people to write every rule or build an accurate model.

  • End-to-end visuomotor policies: in 2016, Levine et al. used a single neural network to map camera images directly to a robot arm's motor torques, completing tasks such as screwing a cap onto a bottle and hanging a coat hanger on a rack.
  • Learning to grasp in the real world: Google used up to 14 robot arms to collect over 800,000 grasp attempts in two months and trained a model that could grasp new objects. QT-Opt went further and applied deep reinforcement learning to vision-based grasping: trained on over 580,000 real grasp attempts, it reached a 96% success rate on unseen objects.
  • Train in simulation, deploy on hardware: OpenAI trained a robot hand with reinforcement learning in simulation and used domain randomization to transfer the policy to a real dexterous hand, which reoriented a block in hand (2018) and solved a Rubik's cube with one hand (2019). ETH Zurich deployed simulation-trained policies on the ANYmal quadruped, achieving agile motion and walking over challenging terrain (2019–2020). In 2021, legged_gym used parallel GPU simulation to cut the training time of a flat-ground walking policy to a few minutes.
  • Embodied simulation platforms and tasks: simulation platforms such as AI2-THOR and Habitat, together with tasks such as point-goal navigation (PointNav), object-goal navigation (ObjectNav), and vision-and-language navigation (VLN), let researchers train and evaluate navigation agents at scale in virtual indoor scenes.

The core shift of this period was from hand-designing every module to letting robots learn from data and interaction. The costs were a sharp rise in the need for data and compute, reward functions that are hard to design, a gap between simulation and reality, and policies that often worked only for the tasks and robots they were trained on.

2.4 Foundation models (2022–present)​

After 2022, progress in large language models and vision-language models quickly reached robotics. Robot capabilities increasingly come from large models pretrained on massive data, which bring language understanding, common sense, and general capabilities across tasks:

  • Planning with language models: SayCan (2022) used a large language model to break down requests such as "I spilled my drink, can you help?" into sequences of skills the robot can execute. Code as Policies (2022) had language models write the code that controls the robot.
  • Robotics Transformers and VLA: RT-1 (2022) trained a Transformer policy on about 130,000 real demonstrations. RT-2 (2023) trained a vision-language model into a vision-language-action model (VLA), enabling the robot to understand instructions that never appeared in its training data.
  • Shared data: Open X-Embodiment (2023) pooled over one million trajectories from 22 robot embodiments to advance training across robots.
  • Low-cost imitation learning: ALOHA with ACT (2023) and Diffusion Policy (2023) showed that tens to hundreds of demonstrations are enough to learn fine-grained bimanual manipulation, greatly lowering the barrier to research.
  • Open source and generalist policies: open models such as OpenVLA (2024) let the community reproduce and fine-tune VLAs. Models such as π0 (2024) and π0.5 (2025) generate continuous actions with flow matching and complete long-horizon tasks such as tidying kitchens and bedrooms in homes they have never seen.
  • World models: models such as Genie, Cosmos, and V-JEPA 2 learn to predict how the environment changes in response to actions, and are used for planning, generating training data, and evaluating policies.

In these systems, the large model usually understands the task and makes decisions, while low-level controllers execute quickly and stably. This is often described as a division of labor between a "brain" and a "cerebellum." Meanwhile, parallel GPU simulation can run thousands of environments at once, new platforms such as humanoid robots keep appearing, and embodied AI has become one of the most closely watched directions in AI research and industry. For the technical evolution of VLA since 2022, see From LLMs to robot policies.

2.5 Lessons from history​

Looking back, several threads run through this history:

  1. The trade-off between modular and end-to-end design never went away. From sense–plan–act versus behavior-based robotics, to modular systems versus end-to-end learning, to today's debate over hierarchical versus monolithic VLA, the core question is the same: how should the system be divided, and what information should pass between the parts?
  2. Capabilities increasingly come from data rather than people. From hand-written rules and hand-built models to learning from interaction and massive data, the scale, quality, and diversity of data increasingly set the ceiling on performance.
  3. Simulation has always been a key tool. Its role has grown from validating algorithms to training reinforcement learning policies in parallel to generating data with world models, and the gap between simulation and reality has persisted throughout.
  4. The body still matters. The same model performs differently on different robots, and cross-embodiment learning remains an active research topic.

The foundations organize learning material for these methods into decision-making, control, perception, and engineering modules.

Summary​

  • In the period of programmed rules, robot capabilities came from rules written by people: first symbolic reasoning and the sense–plan–act paradigm, then behavior-based robotics, which emphasized that intelligence arises from the interaction of body and environment and introduced the idea of embodiment.
  • In the period of model-based control, robots used physical and probabilistic models for estimation, planning, and control; modular systems became reliable, and ROS lowered the barrier to building them.
  • In the period of deep reinforcement learning, robots learned control policies through trial and error; in the period of foundation models, language, vision, and action came together.

Further reading​

References​

  1. Nils J. Nilsson (ed.). Shakey the Robot. SRI International Technical Note 323, 1984.
  2. Rodney A. Brooks. A Robust Layered Control System for a Mobile Robot. IEEE Journal on Robotics and Automation, 1986.
  3. Rodney A. Brooks. Elephants Don't Play Chess. Robotics and Autonomous Systems, 1990.
  4. Rodney A. Brooks. Intelligence without Representation. Artificial Intelligence, 1991.
  5. Tad McGeer. Passive Dynamic Walking. The International Journal of Robotics Research, 1990.
  6. Sebastian Thrun et al. Stanley: The Robot that Won the DARPA Grand Challenge. Journal of Field Robotics, 2006.
  7. Sebastian Thrun, Wolfram Burgard, Dieter Fox. Probabilistic Robotics. MIT Press, 2005.
  8. Sergey Levine et al. End-to-End Training of Deep Visuomotor Policies. JMLR, 2016.
  9. Sergey Levine et al. Learning Hand-Eye Coordination for Robotic Grasping with Deep Learning and Large-Scale Data Collection. 2016.
  10. Dmitry Kalashnikov et al. QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation. 2018.
  11. OpenAI. Learning Dexterous In-Hand Manipulation. 2018; Solving Rubik's Cube with a Robot Hand. 2019.
  12. Jemin Hwangbo et al. Learning Agile and Dynamic Motor Skills for Legged Robots. 2019.
  13. Joonho Lee et al. Learning Quadrupedal Locomotion over Challenging Terrain. 2020.
  14. Nikita Rudin et al. Learning to Walk in Minutes Using Massively Parallel Deep Reinforcement Learning. 2021.
  15. Manolis Savva et al. Habitat: A Platform for Embodied AI Research. 2019.
  16. Michael Ahn et al. Do As I Can, Not As I Say: Grounding Language in Robotic Affordances. 2022.
  17. Jacky Liang et al. Code as Policies: Language Model Programs for Embodied Control. 2022.
  18. Anthony Brohan et al. RT-1: Robotics Transformer for Real-World Control at Scale. 2022.
  19. Anthony Brohan et al. RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control. 2023.
  20. Open X-Embodiment Collaboration. Open X-Embodiment: Robotic Learning Datasets and RT-X Models. 2023.
  21. Tony Z. Zhao et al. Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware. 2023.
  22. Cheng Chi et al. Diffusion Policy: Visuomotor Policy Learning via Action Diffusion. 2023.
  23. Moo Jin Kim et al. OpenVLA: An Open-Source Vision-Language-Action Model. 2024.
  24. Physical Intelligence. π0: A Vision-Language-Action Flow Model for General Robot Control. 2024; π0.5: a Vision-Language-Action Model with Open-World Generalization. 2025.
  25. Jake Bruce et al. Genie: Generative Interactive Environments. 2024.
  26. NVIDIA. Cosmos World Foundation Model Platform for Physical AI. 2025.
  27. Mido Assran et al. V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning. 2025.