Skip to main content

View as Markdown

Page content converted to Markdown. Use the original page link at the end to explore interactive graphics.

5. Key challenges

The previous chapters covered the concepts and history of embodied AI, along with robot skills and body types. This chapter discusses problems that are not yet well solved, which are also the starting point of much research. Knowing these challenges helps you judge what a result actually solves and where it applies.

5.1 Data: scarce, expensive, and hard to reuse​

Large models can draw on huge amounts of text and images from the internet, but robot data must be collected specifically through teleoperation, demonstrations, or autonomous interaction. For example, RT-1's roughly 130,000 demonstrations were collected by 13 robots over 17 months.

The data is also hard to reuse directly: robots differ in sensors, action spaces, and control frequencies, so data from one robot arm may not train another directly. Data quality matters just as much: whether observations and actions are aligned in time, whether action units and coordinate frames are consistent, and whether demonstrations are smooth all affect training results.

Common approaches: pool cross-embodiment datasets such as Open X-Embodiment; generate data at scale in simulation; learn from videos of humans; generate training data with world models; and design more efficient teleoperation and data collection tools.

On this site: Embodied AI datasets, LeRobot course notes.

5.2 The sim-to-real gap​

Training in simulation is safe, cheap, and parallelizable, but simulation always differs from the real world: friction and contact are modeled imperfectly, motors respond with different characteristics and delays, and sensor noise, calibration errors, and timestamp offsets are hard to simulate fully. A policy that performs well in simulation may jitter, slip, or even fall on a real robot.

Common approaches: domain randomization (randomly varying parameters such as mass, friction, and latency during training so that the policy adapts to a range); system identification and actuator modeling (bringing the simulation closer to the real robot); teacher–student training (first training a teacher policy with privileged information available in simulation, then distilling it into a student policy that uses only onboard sensors); light fine-tuning on the real robot; and careful sensor calibration and time synchronization.

On this site: Sensor calibration and sim2real.

5.3 Generalization​

Performing well in the training setting does not mean handling new objects, environments, instructions, or robots. A policy that learned to grasp cups on a lab tabletop may see its success rate drop noticeably when the cup's color, the lighting, or the camera position changes.

Common approaches: collect more diverse data; use the semantic knowledge of pretrained vision-language models, as RT-2 does to understand instructions that never appeared in its training data; train across embodiments; and randomize visual and physical parameters during training.

When evaluating, distinguish reproduction from generalization: testing on the training tasks and environments only shows that the method learned those tasks; testing on new objects and environments is what shows its ability to generalize.

5.4 Long-horizon tasks and compounding errors​

Tasks such as folding clothes or tidying a room take dozens or even hundreds of steps. Small errors at each step accumulate; once the robot reaches a state never seen in the demonstrations, the policy may drift further and further off course. This is called distribution shift.

Common approaches: predict a chunk of actions at a time to reduce the number of decisions, as ACT does with action chunking; break long tasks into subtasks scheduled by a high-level planner; detect failures and attempt recovery; and keep learning from deployment experience.

On this site: ACT action chunking, π0.6 online learning.

5.5 Real-time execution and compute​

Large-model inference may take tens to hundreds of milliseconds, while low-level control cycles usually last only a few to tens of milliseconds. Compute, power, and cooling on a robot are also limited, so large models cannot be deployed freely.

Common approaches: use layered designs in which slower large models make high-level decisions and faster controllers handle low-level execution; use action chunking and asynchronous execution so the robot keeps executing its current plan smoothly while waiting for the next inference; and compress models and accelerate them on the device.

On this site: RTC real-time execution.

5.6 Safety and reliability​

A robot's mistakes have physical consequences: collisions, pinched fingers, falls, or damaged objects. Policies learned from data rarely come with safety guarantees, and their failures are hard to explain. Real applications also require robots to stay stable over long periods of operation, not just to succeed a few times in a demo video.

Common approaches: limit joint velocities, torques, and the workspace; detect collisions and anomalies; keep a conventional controller as a safety fallback; validate thoroughly in simulation before moving to hardware; and during real-robot experiments, make sure someone is supervising and an emergency stop is always within reach.

On this site: the SO-101 + LeRobot hardware tutorial covers hardware connectivity and safety tests step by step.

5.7 Evaluation​

For the same task, different definitions of success, initial conditions, object placements, or numbers of trials can lead to different conclusions. Real-robot evaluation is time-consuming and laborious, and environments are hard to reproduce exactly; reporting only the best run overstates a method's ability.

Common approaches: use public simulation benchmarks such as LIBERO; report experimental conditions, success conditions, and numbers of trials in full; report in-distribution and out-of-distribution results separately; and use world models or simulators to assist evaluation and reduce the time needed on real robots.

5.8 The challenges are connected​

These challenges do not stand alone: real data is scarce, so research turns to simulation, which brings the sim-to-real gap; generalization requires more diverse data; long-horizon tasks test generalization and are limited by real-time constraints; and every improvement needs reliable evaluation to prove it. When you read a new result, ask: which challenge does it mainly address, what does it cost, and under which conditions was it validated?

Summary​

  • Data scarcity, the sim-to-real gap, and generalization are among the most fundamental problems in embodied AI.
  • Long-horizon tasks, real-time execution and compute, and safety and reliability determine whether methods can leave the lab.
  • Reliable evaluation is the prerequisite for judging progress: check success conditions, numbers of trials, and test distributions.

Next steps​

Continue with 6. Technology stack to follow learning modules into the courses. 7. Career roles introduces responsibilities and interview preparation. You can also choose according to your goal:

  • Learn specific principles: go to the foundations.
  • Build a complete system or test a method: choose a project organized into chapters or a standalone experiment in projects.
  • Find papers, data, and code: browse the resources.