Chapters
On this page
6. Technology stack
This chapter introduces the courses that follow. First learn what each technical module does, then study its principles in the foundations and run the methods through projects. Start with an area that interests you and revisit prerequisites as needed.
6.1 Foundation model algorithms
Begin with VLMs, VLAs, and world models. For a task such as “put the red cup on the plate,” a VLM helps interpret the image and instruction, a VLA generates actions from observations, and a world model predicts changes those actions may cause. These capabilities can work together in one system.
VLM: understanding images and language
A vision-language model processes images and text together for description, question answering, target understanding, and task reasoning. Start by learning how images become visual features and how language models process sequences, then study how the two forms of information are fused.
Key modules include Transformers, language models, visual encoders, image-text alignment, and multimodal fusion. For example, a model needs to connect “red cup” in an instruction with an object in the image to provide semantic information for subsequent actions.
Continue learning: Multimodal model foundations → Visual encoders and alignment → Multimodal fusion and VLMs.
VLA: from understanding to action
A vision-language-action model generates actions from information such as images, language instructions, and robot state. Outputs can be control targets for joints or an end effector, or a sequence of continuous actions.
Study action representations, imitation learning, action chunking, diffusion or flow-matching generation, and how model outputs enter the robot's control loop. OpenVLA and π0 illustrate different implementations; post-training explores how interaction outcomes can improve a policy further.
Continue learning: From VLM to VLA → VLA series introduction → OpenVLA / π0.
Practice: π₀.₅ + RECAP reproduction covers data preparation, value models, policy fine-tuning, and simulation evaluation.
World models: predicting the consequences of actions
World models learn how an environment changes over time and in response to actions. In robotics, current observations and candidate actions can be used to predict future states, images, or latent representations. These predictions can support planning, policy learning, or evaluation.
Study action-conditioned prediction, latent dynamics, video prediction, and error over multiple prediction steps. Understand how predictions help select actions and how to check whether the model is trustworthy in new settings.
Continue learning: World model introduction → World model survey and technical approaches. Current coverage focuses on concepts, methods, and research surveys.
A useful starting sequence is “multimodal foundations → VLM → VLA,” followed by world models in the context of planning or policy learning.
6.2 Manipulation algorithms
Manipulation algorithms enable robots to interact with objects. Follow the process from identifying a target to planning a motion, making contact, and completing the task.
- Kinematics and planning: forward kinematics computes end-effector pose from joint angles; inverse kinematics solves for joint configurations. Motion planning adds path and collision constraints. Start with coordinate transformations and kinematics, then study MoveIt 2.
- Imitation learning and action generation: learn policies from observations and actions in demonstrations. ACT predicts action chunks, while Diffusion Policy generates action sequences through a diffusion process. Start with imitation learning, then explore ACT and Diffusion Policy.
- Contact and force control: contact introduces forces and deformation. Study the relationship between position tracking, compliance, and contact forces in impedance and force control.
Practice: the MoveIt 2 simulation project connects planning with execution; ACT bimanual training covers data, training, and evaluation. ACT is an introductory experiment in learned manipulation; the original method does not rely on language instructions.
6.3 Motion control algorithms
Motion control makes joints and the robot body execute actions reliably. Start with error feedback and kinematics, then explore model-based control and policies trained through interaction.
- Feedback control: adjust control signals using the difference between desired and current state. Learn PID / PD, sampling rates, limits, and stability in feedback and PID.
- Model-based control: use dynamics to predict motion and understand how LQR, trajectory tracking, and MPC choose control signals and handle constraints. Follow the controller course chapter by chapter.
- Reinforcement learning control: define observations, actions, and rewards so a policy can learn through repeated interaction. Start with MDPs and PPO, then explore reward design, domain randomization, and sim-to-real transfer.
Practice: Build a quadruped from scratch connects PD control, kinematics, and gaits; MicroDuck RL demonstrates policy training and motion evaluation in simulation.
6.4 Navigation
Understand navigation as “perceive the environment → estimate position → plan a path → execute and avoid obstacles.” Each step depends on consistent coordinate frames and trustworthy sensor data.
- Perception and calibration: learn about cameras, lidar, and IMUs, along with intrinsics, extrinsics, time synchronization, and coordinate projection. Start with sensor calibration and sensor projection.
- Localization and mapping: state estimation combines observations to estimate motion; SLAM estimates the robot's position while building a map. Read odometry, timestamps, and state estimation to understand drift and data alignment.
- Path planning and navigation execution: search for a path in a map and generate executable commands under obstacle and motion constraints. Tools such as Nav2 organize planning, control, and recovery; vision-language navigation also incorporates language goals and semantic information.
Current material covers calibration, state estimation, and tf2. The mapping and navigation project and vision-language navigation project are still planned.
6.5 Simulation engineering
Simulation provides repeatable experiments for algorithms. Start with a minimal scene, then add a robot, sensors, a controller, and a training task.
- Robot and scene descriptions: understand how joints, links, inertia, collision shapes, and actuators form a model through URDF and MJCF.
- Physics and sensor simulation: learn how time steps, contact, friction, and sensor settings affect results, and build environments with MuJoCo or Isaac Sim.
- Training environments and evaluation: organize observations, actions, rewards, and reset conditions into an interface, then study parallel training, domain randomization, and repeated evaluation. Start with the environment interface in Gymnasium.
Practice: MuJoCo simulation introduction creates and runs an environment; MicroDuck RL shows a complete training workflow.
6.6 Data engineering
Robot learning needs sensor observations, actions, and task information organized into trajectories. Data engineering ensures these records can be collected, checked, stored, and used for training correctly.
- Teleoperation and collection: understand how operator inputs map to robot actions while images, joint states, and timestamps are recorded.
- Data organization and quality: understand episode boundaries, task labels, action units, control rates, and calibration metadata, and check for dropped frames, misalignment, and unusual actions.
- Training data preparation: study normalization, training / validation splits, dataset versions, and batch loading, and how the data distribution affects a policy.
Continue learning: Embodied AI datasets explains what to look for in dataset descriptions; the LeRobot course introduces the robot learning toolchain.
Practice: the SO-101 hardware tutorial connects devices, control, and data workflows; ACT bimanual training shows how demonstrations enter training and evaluation.
6.7 Deployment engineering
Deployment connects model outputs to a working robot system, covering the full path from sensor inputs to actuator commands.
- ROS 2 and communication: learn nodes, topics, services, actions, and launch configuration to understand how modules exchange state and commands. Start with ROS 2 engineering basics.
- Model inference and action execution: understand inference latency, control rates, action chunks, and asynchronous execution in RTC real-time execution. Tools such as ONNX Runtime and TensorRT support model execution and inference optimization.
- Control interfaces and debugging: map policy actions to position, velocity, or torque commands, checking units, timing, limits, and abnormal states. Start with controller integration.
Practice: use the MoveIt 2 simulation project to understand system connections, then follow the SO-101 hardware tutorial for device connectivity and a minimal control loop.
6.8 Hardware
Understanding hardware helps you identify what an algorithm can control and observe, and which physical constraints it must respect.
- Mechanics and actuators: learn about links, joints, transmissions, motors, and gearboxes, and understand degrees of freedom, workspace, payload, backlash, and stiffness.
- Electronics and motor drives: learn about power supplies, sensor circuits, encoders, and drives, and the relationship between current, velocity, and position control.
- Embedded systems and communication: learn about MCUs, real-time tasks, and buses such as CAN, including how host commands reach actuators and how state returns.
Start with URDF and robot models, then use the quadruped course to connect modeling with control. Mechanical design and CAN and MCU communication are currently placeholders.
Start with one module
Choose a concrete problem, such as understanding image-text fusion in a VLM, training a manipulation policy, or making a joint track a target. Read the relevant theory, complete an existing experiment, and record its inputs, outputs, and results. Gradually connect individual modules into a complete robot system.
See 7. Career roles for job responsibilities, and the resources for papers, code, and tools.