Single Policy, Many Skills: SMILE Method Advances Whole-Body Motion Generation for Humanoid Robots

A new study published in Aerospace Control and Application presents a learning framework designed to help a humanoid robot acquire diverse whole-body motion skills under a single policy model while maintaining both motion quality and smooth transitions between skills. The research, titled “Whole-body motion strategy intelligent generation method for multi-skilled humanoid robots,” appears in volume 51, issue 2, pages 28–40, with DOI 10.3969/j.issn.1674-1579.2025.02.003. The authors are Zhang Lingjun, Tang Liang, and Liu Lei of the Beijing Institute of Control Engineering and the National Key Laboratory of Space Intelligent Control.

The work addresses a central problem in humanoid robot learning. In many existing systems, a humanoid robot can learn a specific motion skill, such as walking, standing, or squatting, but it often cannot learn many different skills inside one policy model. When a humanoid robot attempts to learn multiple skills together, differences in motion characteristics can create conflicts during policy optimization. Static skills usually require balance and stability, while dynamic skills require the humanoid robot to break static balance briefly, move quickly, and then recover stability. These conflicting demands can make a single policy difficult to train. A second challenge involves transitions between skills. If a humanoid robot focuses only on smooth transitions, individual skill execution may degrade. If it focuses only on skill completion, transitions may become unstable, discontinuous, or inefficient.

To address these issues, the authors propose a method called single model imitation learning for multi-skill efficiency, or SMILE. The method combines goal-conditioned reinforcement learning with generative adversarial imitation learning. It introduces preference-based rewards and a failure-frequency-based priority sampling method. According to the study, SMILE enables a humanoid robot to learn standing, squatting, walking, obstacle jumping, stooping for detailed inspection, object picking, and other human-like whole-body skills. The trained policy achieves smooth transitions across different skills and reaches a success rate of 93.33 percent in simulation. In ablation experiments, removing preference rewards or the failure-frequency-based priority sampling method reduces the success rate to 53.33 percent and 73.33 percent, respectively, while also reducing training efficiency. Compared with a goal-conditioned reinforcement learning method, SMILE alleviates forgetting or performance degradation of previously learned skills during multi-skill training.

1. Why Multi-Skill Learning Remains Difficult for a Humanoid Robot

The paper explains that reinforcement learning has made important progress in humanoid robot skill learning. Prior work has enabled a humanoid robot to walk stably in outdoor environments, perform highly dynamic motions in simulation, and recover from different fallen postures. However, reinforcement learning still has limitations when the desired skills involve large numbers of human-like motion details that are difficult to define explicitly. Imitation learning introduces human demonstration data as prior knowledge and can improve both learning efficiency and human-like behavior. For example, adversarial motion priors have been used to guide a humanoid robot toward straight-knee walking and a more human-like gait.

Even so, many methods focus on one skill or a specific set of motions and are difficult to deploy as a single policy that can handle many skills. Some frameworks separate the upper body and lower body of a humanoid robot and train them with different objectives. This can allow a humanoid robot to imitate diverse human motions with its upper body while keeping the lower body stable, but coordination may be insufficient during highly dynamic motions. Later improvements introduce teacher policies and knowledge distillation to help a deployed policy execute complex dynamic motions such as running and squatting. Other frameworks enable real-time teleoperation of diverse skills for a humanoid robot, but the policy may frequently adjust its gait to maintain stability, making it difficult for the humanoid robot to stand still or perform fine mobile manipulation. Adjustments to the reward function can encourage larger steps or static standing, but such stability-biased strategies may not support high dynamic motion skills.

The study therefore identifies two unresolved problems. The first problem is that different skills have different motion characteristics, and these differences create conflicts in policy optimization. A static skill may require the humanoid robot to maintain static balance, while a dynamic skill may require the humanoid robot to actively break static balance and then recover. The second problem is the trade-off between action completion quality and action continuity during skill transitions. Overemphasizing continuity may reduce the quality of individual skills. Overemphasizing individual skill quality may cause discontinuous or unstable transitions. For a humanoid robot, these issues affect overall task performance and autonomy.

2. BICE-Rob: A Humanoid Robot Platform Built for Space-Oriented Skills

The proposed method is evaluated on BICE-Rob, a humanoid robot system developed for space-related and extraterrestrial exploration scenarios. The humanoid robot system is organized into three layers: a perception layer, a decision layer, and a motion control layer. The perception layer includes sensors such as stereo cameras, depth cameras, lidar, and an inertial measurement unit, along with a perception computer. The decision layer includes a decision computer that transmits perception data to a monitoring terminal, generates task commands autonomously, or processes teleoperation commands. It then distributes commands to the motion control computer. The motion control layer receives commands from the decision layer, solves and generates control commands, and sends them to the joints.

BICE-Rob has a height of 1.50 meters and a weight of 55 kilograms. Its joints are motor-driven, and it has 55 driven degrees of actuation. These include a 13-degree-of-actuation BICE five-finger anthropomorphic dexterous hand with 20 degrees of freedom, 14-degree-of-actuation human-like dual arms, a 2-degree-of-actuation neck, a 1-degree-of-actuation waist, a 3-degree-of-actuation hip, a 1-degree-of-actuation knee, and a 2-degree-of-actuation ankle. The humanoid robot has a highly biomimetic overall configuration. Compared with advanced humanoid robots, BICE-Rob is described as having advantages in meeting both dexterous fine manipulation and highly dynamic motion requirements. Its moderate size, long arm span, high arm degrees of freedom, and dexterous hand degrees of freedom support more human-like fine manipulation. Its hip and knee peak torque of 340 newton-meters gives the humanoid robot strong explosive power and mobility for dynamic motions. BICE-Rob also supports teleoperation for human-in-the-loop fine operations, which is relevant for space station maintenance and extraterrestrial surface exploration.

BICE-Rob Humanoid Robot Attribute Specification
Height 1.50 meters
Weight 55 kilograms
Total driven degrees of actuation 55
Dexterous hand 13-degree-of-actuation BICE five-finger hand, 20 degrees of freedom
Dual arms 14-degree-of-actuation human-like arms
Neck 2 degrees of actuation
Waist 1 degree of actuation
Hip 3 degrees of actuation
Knee 1 degree of actuation
Ankle 2 degrees of actuation
Peak hip and knee torque 340 newton-meters
Teleoperation capability Supported for human-in-the-loop fine operations

The task considered in the paper is for the humanoid robot to learn ten typical skills required for on-orbit maintenance and deep space exploration. These skills are selected to reflect basic needs for autonomy and self-preservation. The humanoid robot must learn the skills under a single policy model using a multi-skill whole-body motion dataset based on human demonstrations. It must also transition coherently and smoothly among these skills. The ten skills are standing, squatting, walking, carrying, close inspection, long-duration cruising, multi-directional picking, sit-to-stand transition, jumping over obstacles, and high-dynamic striding.

Number Skill Application Scenario
1 Standing Maintain balance in an upright posture while completing coordinated upper-limb motion or dexterous operation
2 Squatting Pick up samples from an extraterrestrial surface or maintain equipment interfaces at a low position
3 Walking Transition among forward, backward, and lateral walking for autonomous charging, emergency stopping, or obstacle avoidance
4 Carrying Transfer supply boxes in a space station or carry equipment and construction materials on an extraterrestrial surface
5 Close inspection Bend or stoop to observe instruments inside a cabin or inspect the surrounding environment
6 Long-duration cruising Inspect between orbital modules or conduct long-distance inspection and exploration on an extraterrestrial planet
7 Multi-directional picking Pick up objects in different directions and grasp handrails or fixed points at different orientations
8 Sit-to-stand transition Change posture in front of a control console or inside a cockpit, or shift from different standing postures to sitting to save energy
9 Jumping over obstacles Quickly jump over trenches, pits, or other obstacles to improve mobility and obstacle-crossing ability
10 High-dynamic striding Quickly switch support points and rapidly cross obstacles or uneven terrain

3. The SMILE Method: One Model for Multiple Humanoid Robot Skills

The SMILE method is designed as a single-model multi-skill imitation learning approach. Its core components include a discriminator, a policy network, a value network, a reference motion sampling method, and a reward shaping method. SMILE trains by alternating optimization between the discriminator and the policy network. The discriminator distinguishes between trajectories generated by the policy and reference trajectories. It continuously updates and improves its ability to discriminate different motion styles. At the same time, the policy is optimized under the combined guidance of imitation rewards, discriminator rewards, preference rewards, and physical constraint rewards. The policy attempts to “fool” the discriminator so that the generated motions become more human-like. This adversarial training process encourages the policy to master diverse human-like skills.

3.1 Human Motion Retargeting for a Humanoid Robot

To transfer human demonstration data to the humanoid robot, the method addresses differences in skeletal topology, driven degrees of freedom, and joint axis definitions between humans and BICE-Rob. The retargeting process has two main steps. First, the shape parameters of the SMPL human parametric model are optimized so that, under the same reference pose, joint spatial positions are as close as possible to the corresponding joint positions of BICE-Rob. This reduces motion distortion caused by differences between the human model and the humanoid robot’s body proportions. Second, after obtaining an SMPL model matched to the humanoid robot’s body dimensions, selected joint positions are used as key points to establish a spatial position matching relationship between BICE-Rob and the human motion data. Inverse kinematics is then used to solve for the humanoid robot’s joint angles so that BICE-Rob’s spatial trajectory aligns with the human demonstration.

3.2 Network Architecture and Training

The policy network is a six-layer multilayer perceptron with 2048, 1536, 1024, 1024, 512, and 512 neurons. The hidden-layer activation function is SiLU, and the output layer uses a linear mapping. The value network uses the same six-layer multilayer perceptron structure. The discriminator network contains two hidden layers with 1024 and 512 neurons, respectively, and uses ReLU as the hidden-layer activation function. The discriminator loss function is designed and calculated according to the adversarial motion priors framework, and the discriminator network parameters are updated based on this loss. The policy network is trained using proximal policy optimization.

3.3 State Space and Action Space

The proprioceptive state of the humanoid robot is a 298-dimensional vector representing the current actual motion state, including all joint poses, 3D positions, angular velocities, and linear velocities. To improve training stability, poses are represented in a 6D form. The goal state is a 480-dimensional vector used to measure the difference between actual motion and reference motion. It includes the reference pose, position, linear velocity, angular velocity, relative rotation error between reference and current pose, and reference position at the next time step. The action space is a 19-dimensional vector defined as target joint positions for a proportional-derivative controller. The study does not consider the driven degrees of freedom of the humanoid robot’s neck, dexterous hands, wrists, or ankle inversion and eversion.

3.4 Reward Function and Preference Rewards

The reward function includes four parts: imitation reward, discriminator reward, preference reward, and physical constraint reward. The imitation reward, discriminator reward, and physical constraint reward are generally consistent with prior work. The imitation reward includes six subterms: joint angle imitation, joint angular velocity imitation, position imitation, orientation imitation, velocity imitation, and angular velocity imitation. Each subterm is computed using an exponential decay function. The discriminator reward is calculated from the discriminator output. When the state distribution generated by the humanoid robot policy is closer to the reference motion state distribution, the discriminator gives a higher reward. The physical constraint reward encourages the humanoid robot to complete the task with lower energy consumption and to reduce high-frequency foot jitter.

The paper introduces preference rewards to constrain behaviors that the discriminator may accept but humans would consider unreasonable. Preference rewards are reward signals that reflect human subjective preferences and directly describe behavior features that humans consider good or bad. They include a tripping penalty, a sliding penalty, and a foot orientation reward. The tripping penalty penalizes signs of tripping, such as the toe striking the ground and causing the body to lean forward, or actual falls. It encourages the humanoid robot to avoid severe loss of balance or falling. The sliding penalty penalizes obvious sliding of the foot relative to the ground when the foot is in contact and supporting the body. It encourages stable foot contact and a more human-like gait by avoiding friction-based foot sliding. The foot orientation reward penalizes non-vertical foot orientation. It encourages the humanoid robot to keep the sole flat against the ground so that the normal direction of the sole is as close as possible to the gravity direction. This helps avoid toe standing, single-edge foot contact, and similar undesirable postures.

3.5 Initialization and Termination Conditions

The initial state distribution and termination conditions strongly affect a humanoid robot’s ability to learn diverse skills. The method uses reference state initialization, in which a state is randomly sampled from a reference motion clip as the initial state for each episode. This reduces the difficulty of policy exploration and improves learning efficiency. During training, an early termination mechanism is introduced. When the average deviation between the humanoid robot’s joint positions and the reference motion exceeds 0.25 meters, the current episode is terminated.

3.6 Failure-Frequency-Based Priority Sampling

To avoid catastrophic forgetting, severe fluctuations in the learning curve, and difficulty converging when a single policy model learns multiple skills, the study applies the idea of prioritized experience replay from reinforcement learning to imitation learning. The proposed failure-frequency-based priority sampling method automatically evaluates task success during training and prioritizes reference motion samples that the policy struggles to learn. Initially, the policy samples all reference motions uniformly. When the policy repeatedly fails on certain samples, the sampling probability of those difficult samples is increased. Samples that have already been mastered are assigned a smaller sampling probability to consolidate existing skills and avoid complete forgetting.

The sampling probability of a reference motion is calculated from accumulated failure counts with a smoothing term and a power parameter. The failure count for a sample is updated with a decay coefficient. If a sample fails during an evaluation, its failure count increases; otherwise, it decreases according to the decay term. This adaptive sampling strategy allows the humanoid robot to focus more training effort on difficult skills while still revisiting learned skills.

4. Simulation Results: A Humanoid Robot Learns Ten Whole-Body Skills

Simulation experiments are conducted in Isaac Gym. The policy runs at 30 hertz, and the simulation runs at 60 hertz. Training uses an Intel i9-13900K CPU at 5.80 gigahertz and an NVIDIA RTX 4090 GPU. To evaluate SMILE, the researchers select 15 representative groups of diverse human demonstration data from the AMASS dataset. These include stable standing, squatting, carrying, multi-directional picking, stooping for inspection, forward walking, lateral-to-forward walking transition, backward-to-forward walking transition, long-duration cruising, stepping in place, jumping over obstacles, aerial turning, standing-to-high-dynamic-striding transition, walking-to-high-dynamic-striding transition, and sit-to-stand transition.

The trained BICE-Rob humanoid robot demonstrates diverse motion skills. In multi-directional object picking, the humanoid robot coordinates its whole body, including waist rotation, leg adjustment, and foot changes, to achieve posture adjustment and turning in different directions. This shows good whole-body coordination. In forward and backward walking transitions, the humanoid robot shows a human-like heel-to-toe gait. The front foot’s heel lands first, the center of gravity shifts, and the rear foot leaves the ground from heel to toe. When transitioning quickly from forward walking to a stop, the humanoid robot uses both toes touching the ground to stop suddenly, then shifts to backward walking. The backward walking shows a human-like pattern in which the front foot lands on the heel and the rear foot changes from toe contact to heel contact.

In lateral-to-forward walking, the humanoid robot initially faces forward and gradually completes the transition from lateral gait to forward gait through coordinated whole-body motion. During the turn, one foot is flat on the ground while the other foot uses toe contact to rotate the body. This turning method is consistent with common human turning characteristics. In stooping for detailed inspection, because BICE-Rob has only one rotational degree of freedom at the waist, the stooping motion is achieved by coordinating whole-body joints, mainly through hip flexion to tilt the torso forward. This posture demonstrates coordinated control under limited degrees of freedom. In high-dynamic striding, the humanoid robot coordinates its whole body to achieve the desired leg lift height. The support leg uses toe contact, while the other leg lifts rapidly. During the motion, the humanoid robot briefly has both feet off the ground, then the support leg touches down again and continues the stride. This reflects whole-body coordination, dynamic stability, and balance control.

In obstacle jumping, the humanoid robot successfully achieves a moment in which both feet leave the ground. In a series of static motions, the humanoid robot performs different standing poses, squatting, and sitting oriented in different directions. The humanoid robot can also achieve a distinct human-like sitting posture in which the upper body leans backward to maintain stability. This posture places higher demands on whole-body balance control and demonstrates the humanoid robot’s ability to coordinate whole-body motion and maintain posture in complex static poses.

5. Ablation Study: Preference Rewards and Priority Sampling Matter

To verify the influence of core components in SMILE, the researchers conduct ablation experiments on the reference motion sampling method and the preference reward function. The learning curves show clear differences. The failure-frequency-based priority sampling method has notable advantages in improving success rate and accelerating training efficiency when the humanoid robot learns diverse skills. With SMILE, the policy rapidly reaches a success rate of about 93.33 percent during training. With uniform sampling, the policy reaches a maximum of only 73.33 percent, a difference of 20 percentage points. From the perspective of average episode length, SMILE produces longer episodes during training and faster growth. The total reward trend further verifies the advantage of priority sampling. The reward under SMILE increases significantly faster than under uniform sampling and stabilizes at a higher level in later training.

Method Maximum Success Rate Maximum Average Episode Length Training Efficiency Observation
SMILE 93.33 percent 309 steps Reaches 73.33 percent at 1,500 evaluations, 86.66 percent at 3,000, and 93.33 percent at 4,500
Without preference rewards 53.33 percent 177 steps Reaches 33.33 percent at 1,500 evaluations and only reaches its maximum at 16,500 evaluations
Without priority sampling 73.33 percent Not higher than SMILE in the reported comparison Lower reward growth and lower final reward than SMILE

The ablation results for preference rewards show three main advantages. First, preference rewards significantly improve task success rate. SMILE reaches about 93.33 percent success, compared with a maximum of 53.33 percent when preference rewards are removed, an improvement of 40 percentage points. In terms of average episode length, SMILE reaches up to 309 steps, while the method without preference rewards reaches only 177 steps. This suggests that preference rewards reduce the probability of early failure and allow the humanoid robot to explore and exploit a larger state-action space. The average reward curve also shows that SMILE maintains a higher average reward, indicating that preference rewards work together with imitation rewards, discriminator rewards, and physical constraint rewards to guide policy optimization.

Second, preference rewards improve training efficiency and stability. The learning curves for success rate, average episode length, and average reward show that SMILE grows faster and more stably. At 1,500 evaluations, SMILE reaches 73.33 percent success, 86.66 percent at 3,000 evaluations, and 93.33 percent at 4,500 evaluations. In contrast, the method without preference rewards reaches only 33.33 percent at 1,500 evaluations and does not reach its maximum 53.33 percent until 16,500 evaluations. The average episode length and average reward under SMILE also grow faster and more stably, without obvious drops or setbacks. The method without preference rewards shows clear fluctuations near 1,500 evaluations. This indicates that preference rewards reduce policy performance fluctuations during training, reduce the probability of early failure or local optima, and significantly improve training efficiency and stability.

Third, preference rewards improve the human-like performance of the multi-skill policy. The discriminator reward curve shows that SMILE quickly exceeds 1.0 at about 580 steps during early training and steadily rises to about 1.6 by the end of training. In contrast, the method without preference rewards grows slowly and reaches only about 0.6. The discriminator reward directly reflects the similarity between the generated policy state distribution and the human demonstration state distribution. A higher discriminator reward means the policy generates motions closer to human behavior patterns, with more obvious human-like characteristics. Without preference rewards, the method lacks clear human preference guidance, so the generated motions differ more from human behavior patterns, limiting human-like characteristics and overall performance.

6. Comparison with Goal-Conditioned Reinforcement Learning

The study also compares SMILE with a goal-conditioned reinforcement learning method, OmniH2O, and with an OmniH2O policy fine-tuned using the SMILE reward function. The training procedure has two stages. In the first stage, BICE-Rob is trained to learn diverse motion skills using the reward function and hyperparameters of OmniH2O, producing an initial policy. In the second stage, the reward function is adjusted to match the imitation reward, constraint reward, and preference reward used in SMILE, but without the discriminator reward. This fine-tuning stage examines whether a reinforcement learning method can learn more complex, diverse, and dynamic motion skills when the humanoid robot already has basic skills.

The initial policy shows a slow upward trend in average episode length and reward. In later training, the policy gradually stabilizes, but the final average episode length and reward remain low. This indicates that the initial policy learns only a small number of skills. After the second-stage fine-tuning, the learning curve differs. In early training, episode length and reward rise quickly and clearly exceed the initial policy. This suggests that fine-tuning based on the initial policy can rapidly improve the humanoid robot’s imitation performance for diverse skills. However, the second-stage learning curve fluctuates more strongly. This phenomenon indicates that without the discriminator reward to constrain motion performance, the policy’s ability to learn diverse skills is limited.

Visualization of the policy generated by the goal-conditioned reinforcement learning method shows that BICE-Rob adopts a conservative “standing-first” style in the initial policy. The humanoid robot has some static skill imitation ability, prioritizing lower-limb stability and minimizing lateral or large center-of-mass shifts. This strategy avoids small “stomping” adjustments and makes overall motion smoother. However, it has several limitations. First, it overemphasizes standing stability, so imitation performance is poor for lateral movement, turning, or large dynamic motions. Second, lower-limb motion is relatively rigid and cannot easily adapt to actions requiring larger steps or rapid foot adjustments, such as lateral walking, obstacle jumping, and multi-directional object picking. Third, action completion is low; the humanoid robot often substitutes waist rotation or large arm swings for leg movement, causing imitation failure.

After second-stage fine-tuning, BICE-Rob can attempt to imitate more diverse action sequences. However, because there is no discriminator reward to explicitly constrain motion imitation quality, the fine-tuned policy has several problems. First, action completion decreases noticeably. When demonstrating dynamic skills, BICE-Rob struggles to accurately reproduce key motion features and tends to use small steps or body swaying instead. For example, obstacle jumping should include a key feature in which both feet leave the ground at the same time, but BICE-Rob replaces the airborne phase with continuous small steps. This “motion deformation” indicates that without the effective constraint of the discriminator reward, the humanoid robot’s action completion can decline significantly.

Second, action execution efficiency decreases significantly. When performing tasks involving larger strides or rapid movement, such as walking, high-dynamic striding, long-duration cruising, and stooping for inspection, BICE-Rob often takes too-small steps and frequently adjusts its feet. This leads to unsmooth motion transitions and slow speed, severely limiting the humanoid robot’s overall performance in multi-skill tasks. Third, previously learned actions degrade or are forgotten. Because the reward function reduces the reward for “standing-first” and large-step movement during fine-tuning, BICE-Rob shows obvious degradation when performing actions that the initial policy had already learned well. For example, forward and backward walking with large strides in the initial policy degrades into inefficient small-step shuffling after fine-tuning. This phenomenon indicates that although the humanoid robot acquires new skills, it often implements them in nonstandard or inefficient ways, causing the overall policy to trade one capability for another.

Comparison Aspect SMILE Goal-Conditioned Reinforcement Learning and Fine-Tuning
Core guidance Imitation reward, discriminator reward, preference reward, physical constraint reward Reward adjustment and policy fine-tuning, with no discriminator reward in the second stage
Initial behavior Learns multiple skills under one model with smooth transitions Initial policy shows standing-first conservative style and limited skills
Dynamic motion completion Humanoid robot can achieve skills such as obstacle jumping and high-dynamic striding Humanoid robot substitutes small steps or body swaying for key dynamic features
Execution efficiency Higher efficiency and smoother transitions reported under SMILE Small steps, frequent foot adjustment, slow and unsmooth motion
Forgetting or degradation Alleviates forgetting or performance degradation of original skills Previously learned large-stride walking degrades after fine-tuning

7. What the Results Mean for Humanoid Robot Autonomy in Space

The study frames its motivation around space environments. Compared with planetary rovers and space robotic arms, a humanoid robot has unique advantages in generality, adaptability, and dexterity. Inside a space station, a humanoid robot can directly use tools, equipment, and interfaces designed for astronauts without additional adaptation, reducing the cost of spacecraft design and maintenance. On extraterrestrial surfaces, a humanoid robot with a human-like body structure and highly flexible whole-body motion control can learn more diverse whole-body and mobile manipulation skills. This helps it adapt to complex, unknown, and unstructured environments, improving the efficiency and autonomy of exploration tasks. To support space station use and maintenance, lunar research station construction, and deep space exploration, a humanoid robot must master diverse skills and transition coherently and smoothly among them.

The paper states that SMILE provides a new approach for humanoid robot multi-skill imitation learning. It effectively alleviates conflicts caused by different skill characteristics during policy optimization and balances transition motion quality and continuity. Compared with a policy generation method without preference rewards, SMILE significantly improves task success rate through preference reward guidance, improves training efficiency and stability, and enhances the human-like performance of the policy. The authors suggest that SMILE can lay a technical foundation for deploying whole-body motion strategies for multi-skilled humanoid robots in future space station and extraterrestrial exploration missions.

Future work will focus on sim-to-real transfer. The researchers plan to use a teacher-student paradigm to further improve the adaptability and engineering feasibility of the proposed method on the real BICE-Rob humanoid robot platform. The authors also note that SMILE has good scalability. Although the method is validated on a relatively limited skill set, similar policy generation methods can be used for larger or more diverse datasets. Because preference rewards and failure-frequency-based priority sampling can guide learning across different skills, they are expected to maintain high learning efficiency and policy performance even when the number of skills increases. Future research will also explore how a humanoid robot can handle continuously increasing new skill learning demands while preserving existing skills when conflicts between skill characteristics become more significant, further improving the scalability and adaptability of policy generation methods.

8. Key Contributions of the SMILE Method for Humanoid Robot Learning

  • The study identifies two central challenges for multi-skill humanoid robot learning: conflicting motion characteristics across skills, and the trade-off between skill completion quality and transition continuity.
  • It proposes SMILE, a single-model multi-skill efficient imitation learning method that combines goal-conditioned reinforcement learning with generative adversarial imitation learning.
  • It introduces preference rewards to guide the humanoid robot away from motions that a discriminator may accept but humans would consider unreasonable, including tripping, foot sliding, and non-vertical foot orientation.
  • It applies a failure-frequency-based priority sampling method that automatically increases the sampling probability of reference motions that the humanoid robot repeatedly fails to learn.
  • The trained humanoid robot learns standing, squatting, walking, carrying, close inspection, long-duration cruising, multi-directional picking, sit-to-stand transition, obstacle jumping, and high-dynamic striding.
  • Simulation results show a 93.33 percent success rate for SMILE, compared with 53.33 percent after removing preference rewards and 73.33 percent after removing priority sampling.
  • Compared with goal-conditioned reinforcement learning and reward-based fine-tuning, SMILE alleviates forgetting or performance degradation and improves action completion, execution efficiency, and human-like behavior.

The research therefore presents a structured path toward a humanoid robot that can learn many whole-body skills in one policy model, transfer among them smoothly, and retain performance across different motion types. For a humanoid robot intended for space station maintenance and extraterrestrial exploration, such capabilities are directly connected to autonomy, environmental adaptability, and self-preservation. The SMILE method does not solve every sim-to-real problem, and the authors explicitly identify real-platform deployment as future work. Nevertheless, the reported simulation results and ablation studies provide quantitative evidence that preference rewards and failure-based priority sampling are important components for multi-skill humanoid robot policy generation. The work also supports the broader trend of using imitation learning and reinforcement learning together so that a humanoid robot can acquire human-like motion details while still optimizing task success and stability.

The publication appears in Aerospace Control and Application, 2025, volume 51, issue 2, pages 28–40, under DOI 10.3969/j.issn.1674-1579.2025.02.003. The English citation is listed as Zhang L J, Tang L, Liu L. Whole-body motion strategy intelligent generation method for multi-skilled humanoid robots. Aerospace Control and Application, 2025, 51(2): 28–40, in Chinese.

Scroll to Top