My research focuses on enabling humanoid robots to learn basketball shooting skills through deep reinforcement learning (DRL) algorithms. The motivation for this work stems from the rapid advancement of humanoid robotics technology and the increasing demand for autonomous decision-making capabilities in complex, dynamic environments. Unlike traditional control methods that depend heavily on precise mathematical modeling and manually designed rules, DRL offers a data-driven paradigm that allows robots to acquire complex skills through trial-and-error interactions with their environments. In this thesis, I address two primary challenges in applying DRL to humanoid robot basketball shooting: sample inefficiency and slow convergence during training, as well as the limitations of discrete action space modeling in dynamic scenarios and the sparse reward problem. To tackle these challenges, I propose two novel algorithm frameworks: the PSR-SAC algorithm (Priority Success Recent Experience Replay-SAC) and the H-PPO algorithm (Hybrid Proximal Policy Optimization), both specifically designed for the humanoid robot basketball shooting task.
The remainder of this article is organized as follows. First, I provide a comprehensive review of the background and related work in humanoid robot motion control, basketball shooting methods, and improvements to classical reinforcement learning algorithms. Next, I present the theoretical foundations of deep learning and reinforcement learning, including convolutional neural networks, attention mechanisms, and the SAC and PPO algorithms. Then, I detail the system design and experimental methodology. Finally, I present experimental results in simulated environments, followed by conclusions and future outlook.

Background and Motivation
Humanoid robots have evolved from conventional industrial manipulators into versatile platforms capable of operating in human-centric environments. The unique advantage of humanoid robots lies in their human-like morphology, enabling them to interact with environments designed for humans. Applications such as disaster response, domestic assistance, education, and public service have driven the need for more advanced autonomous control strategies. In 2023, the Chinese Ministry of Industry and Information Technology released guidelines emphasizing the transformative potential of humanoid robots as disruptive products following computers, smartphones, and electric vehicles. However, realizing the full potential of humanoid robots requires solving high-dimensional control problems, dynamic balance maintenance, and perceptual decision-making. Among these challenges, acquiring complex manipulation skills such as basketball shooting is particularly demanding, as it requires precise motor control, real-time perception, and adaptive decision-making.
Traditional approaches to humanoid robot control rely on classical methods such as PID control, zero moment point (ZMP) stabilization, model predictive control (MPC), and inverse kinematics. While these methods have served as the foundation of humanoid robotics, they exhibit significant limitations in complex, dynamic environments. These limitations include strong dependency on accurate mathematical models, poor adaptability to unstructured scenarios, and high development costs associated with manual rule design. The emergence of DRL has introduced a fundamentally different paradigm. By framing control problems as Markov Decision Processes (MDP) and enabling robots to learn optimal policies via reward feedback, DRL allows humanoids to develop skills through continuous interaction with the environment. This approach reduces reliance on precise models and handcrafted rules while providing a natural mechanism for adaptation to uncertain and dynamic situations.
Despite its promise, applying DRL to humanoid robots for basketball shooting presents several obstacles. First, the high-dimensional visual input combined with the continuous control signals of the humanoid produces a large and complex policy space, resulting in sample inefficiency and slow convergence. Second, the basketball shooting task involves discrete action choices (e.g., turn left, turn right, walk forward, shoot) as well as continuous parameters (e.g., movement duration, shooting strength), which makes action space modeling challenging. Third, the reward signal in shooting tasks is inherently sparse; the robot only receives positive feedback when the ball successfully enters the basket, which may require many episodes of exploration. These issues form the central focus of my research.
Review of Related Methods
Research on humanoid robot control methods can be broadly categorized into two streams: classical control theory and learning-based methods. Classical methods, including PID control, ZMP-based walking, MPC, and optimization-based approaches, have been extensively explored. For example, predictive PID controllers simplify gait generation for bipedal robots, while nonlinear MPC frameworks enable robust disturbance rejection by integrating ankle, hip, and step adjustment strategies. Inverse kinematics and LQR techniques provide precise joint-level control, and biologically inspired methods such as central pattern generators emulate neural mechanisms for locomotion. However, these approaches share a common weakness: they require accurate analytical models and are sensitive to modeling errors. In dynamic environments with variable object positions or unexpected perturbations, their adaptability and robustness are limited.
Learning-based methods, particularly DRL, have demonstrated remarkable success in humanoid robot control. Applications include gait generation, push recovery, soccer skills, object manipulation, and navigation in complex terrains. For instance, researchers have applied PPO variants to achieve end-to-end gait control on uneven terrain, used DQN for push-recovery control, and employed hierarchical DRL frameworks for multi-skill locomotion. More recently, large-scale simulated training with transformer-based policies has enabled humanoid robots to perform zero-shot sim-to-real transfer on challenging terrains. These studies highlight the potential of DRL in robotics, yet most approaches either focus on continuous action spaces or discrete action spaces separately, leaving a gap for tasks that naturally require hybrid action spaces.
Focusing specifically on basketball shooting, early research used fuzzy logic control or Petri-net-based architectures to coordinate perception and action. These methods achieved reasonable performance in static scenarios but struggled with dynamic environments. More recent work has applied Q-learning or DQN to enable robots to learn shooting skills through trial and error. However, these attempts were limited by fixed shooting strengths and discrete action primitives, leading to low success rates and slow convergence when the ball position varied. The work I build upon in this thesis includes a DQN-based end-to-end framework for humanoid basketball shooting, which demonstrated the feasibility of DRL for this task but still suffered from the aforementioned challenges.
Regarding improvements to reinforcement learning algorithms, several important directions have been established. Prioritized experience replay (PER) uses TD-error as a priority metric to sample high-value experiences more frequently, improving sample efficiency. Other strategies include distribution correction approaches, mixed experience replay with successful episodes, and methods emphasizing recent transitions. In the realm of hybrid action spaces, research has explored parameterized action MDPs where each discrete action is associated with continuous parameters. Algorithms such as PA-DDPG, P-DQN, and MP-DQN have been developed to handle discrete-continuous action spaces, although they often optimize the discrete and continuous components independently or require complex Q-function designs.
Preliminaries and Theoretical Foundations
In this section, I briefly review the fundamental concepts that underpin the proposed methods. Deep learning, particularly convolutional neural networks (CNNs) and attention mechanisms, provides the representational backbone for perception. Reinforcement learning supplies the decision-making framework, with specific algorithms such as SAC and PPO serving as the base learners.
Convolutional Neural Networks
CNNs are specialized neural architectures designed for grid-like data, such as images. Their core components include convolutional layers, pooling layers, activation functions, and fully connected layers. Convolution layers apply learnable filters to local regions of the input, capturing hierarchical features from edges to object parts. The operation of a single convolutional layer is expressed as:
\[
Y(i,j) = \sum_{c}\sum_{m}\sum_{n} X(i+m, j+n, c) \cdot W(m,n,c) + b
\]
where \(X\) is the input tensor, \(W\) denotes the convolutional kernel, \(b\) is the bias, and \(Y\) is the output feature map. The output spatial dimension can be computed as:
\[
H_{out} = \frac{H_{in} + 2P – (H_{k} – 1) \cdot D – 1}{S} + 1
\]
where \(H_{in}\) and \(H_{out}\) are input and output heights, \(P\) is padding, \(D\) is dilation, and \(S\) is stride.
Attention Mechanisms and CBAM
Attention mechanisms allow neural networks to focus on task-relevant parts of the input while suppressing irrelevant information. In the context of visual representation learning, attention can be applied along channel and spatial dimensions. The Convolutional Block Attention Module (CBAM) combines these two types of attention sequentially. Given an intermediate feature map \(F \in \mathbb{R}^{C \times H \times W}\), CBAM computes:
\[
F’ = M_{c}(F) \otimes F, \quad F” = M_{s}(F’) \otimes F’
\]
where \(M_{c}\) is the channel attention map and \(M_{s}\) is the spatial attention map. The channel attention is computed via:
\[
M_{c}(F) = \sigma(\text{MLP}(\text{AvgPool}(F)) + \text{MLP}(\text{MaxPool}(F)))
\]
where \(\sigma\) denotes the sigmoid activation. Spatial attention is formulated as:
\[
M_{s}(F’) = \sigma(f^{7 \times 7}([\text{AvgPool}(F’); \text{MaxPool}(F’)]))
\]
where \(f^{7 \times 7}\) represents a \(7 \times 7\) convolution operation.
Markov Decision Processes and Reinforcement Learning
Reinforcement learning problems are typically formalized as Markov Decision Processes defined by the tuple \((S, A, P, R, \gamma)\). At each time step \(t\), the agent in state \(s_t\) selects an action \(a_t\), receives a reward \(r_t\), and transitions to the next state \(s_{t+1}\) according to the transition probability \(P(s_{t+1} | s_t, a_t)\). The objective is to maximize the expected cumulative discounted reward:
\[
G_t = \sum_{k=0}^{\infty} \gamma^{k} r_{t+k}
\]
The state value function and action value function under a policy \(\pi\) are defined respectively as:
\[
V^{\pi}(s) = \mathbb{E}_{\pi}[G_t | s_t = s]
\]
\[
Q^{\pi}(s,a) = \mathbb{E}_{\pi}[G_t | s_t = s, a_t = a]
\]
The Bellman equations provide recursive relationships for these value functions, and the goal of reinforcement learning is to find an optimal policy \(\pi^{*}\) that maximizes the expected return.
SAC Algorithm
Soft Actor-Critic (SAC) is an off-policy maximum entropy deep reinforcement learning algorithm. It augments the standard objective with an entropy regularization term to encourage exploration:
\[
J(\pi) = \mathbb{E}_{\pi}\left[\sum_{t} \gamma^{t}\left(r_t + \alpha \mathcal{H}(\pi(\cdot | s_t))\right)\right]
\]
where \(\mathcal{H}\) denotes the Shannon entropy and \(\alpha\) is the temperature parameter balancing exploration and exploitation. The Critic network is trained by minimizing a soft Bellman residual:
\[
J_{Q}(\theta) = \mathbb{E}_{(s,a,r,s’)\sim \mathcal{D}} \left[\left(Q_{\theta}(s,a) – (r + \gamma \hat{Q}_{\bar{\theta}}(s’, a’) – \alpha \log \pi_{\phi}(a’|s’))\right)^2\right]
\]
The Actor network maximizes the expected Q-value plus entropy:
\[
J_{\pi}(\phi) = \mathbb{E}_{s \sim \mathcal{D}} \left[\mathbb{E}_{a \sim \pi_{\phi}}[\alpha \log \pi_{\phi}(a|s) – Q_{\theta}(s,a)]\right]
\]
SAC also automatically tunes the temperature parameter by minimizing:
\[
J(\alpha) = \mathbb{E}_{s \sim \mathcal{D}}[-\alpha \log \pi_{\phi}(a|s) – \alpha \bar{H}]
\]
where \(\bar{H}\) is a target entropy hyperparameter.
PPO Algorithm
Proximal Policy Optimization (PPO) is an on-policy algorithm that improves training stability by constraining the policy update magnitude. The clipped surrogate objective is:
\[
L^{CLIP}(\theta) = \mathbb{E}_{t}\left[\min\left(r_t(\theta) \hat{A}_t, \text{clip}(r_t(\theta), 1-\epsilon, 1+\epsilon) \hat{A}_t\right)\right]
\]
where \(r_t(\theta) = \frac{\pi_{\theta}(a_t|s_t)}{\pi_{\theta_{old}}(a_t|s_t)}\) is the importance sampling ratio and \(\epsilon\) is the clipping parameter. PPO achieves a balance between sample efficiency and implementation simplicity, making it a reliable baseline for many control tasks.
Simulation Environment and Robot Platform
I conducted all experiments in the Webots simulation environment, an open-source robotics simulator that uses the ODE physics engine and OpenGL rendering. Webots supports multiple programming languages, including C++ and Python, and provides a rich set of robot models and sensors. The Robotis OP2 humanoid robot, a 20-degrees-of-freedom open-source platform, was chosen as the experimental platform. It features a head camera, joint position sensors, and gyroscope/accelerometer sensors, making it suitable for vision-based reinforcement learning research.
In the simulation, the robot’s control architecture relies on two core modules: Gait Manager and Motion Manager. Gait Manager provides walking and turning behaviors through configurable parameters, and Motion Manager handles complex action sequences such as the predefined shooting motion. For finer-grained control, I also directly commanded individual joint motors.
The experimental scenario was constructed following the rules of the FIRA HuroCup basketball competition. The robot starts at a fixed position, and the ball is placed on a pole somewhere within a designated arc area. The basket (a red cylindrical target) is located at the origin of the coordinate system. Two scenarios were considered: a fixed ball position and a random ball position, the latter of which introduces additional difficulty due to the variability in the robot-to-ball and robot-to-basket distances.
Method 1: PSR-SAC for Humanoid Robot Basketball Shooting
Problem Analysis
Traditional approaches to humanoid robot basketball shooting heavily rely on hand-coded visual positioning and fuzzy control, which are time-consuming to design and poorly adapted to dynamic environments. Using DRL to replace these hand-crafted rules offers a more scalable and autonomous solution. However, standard DRL algorithms suffer from two well-known weaknesses: low sample utilization and slow convergence. These issues are particularly acute in tasks with high-dimensional visual inputs and sparse rewards, such as basketball shooting. I therefore set out to address both aspects: improving the visual representation learning and improving the experience replay mechanism.
Subtask Decomposition
The complete basketball shooting task was decomposed into three sequential subtasks: (1) approaching the ball, (2) picking up the ball, and (3) shooting the ball into the basket. This decomposition reduces the complexity of each learned policy and improves training efficiency. The following table summarizes the three subtasks and their objectives:
| Subtask | Objective | Termination Condition |
|---|---|---|
| Subtask 1: Approach Ball | Move to a position in front of the ball | Distance between robot and ball < 0.12 m |
| Subtask 2: Pick Up Ball | Retrieve the ball using arm motions | Ball is located above the robot’s hand |
| Subtask 3: Shoot Basket | Throw the ball into the basket | Ball overlaps with the basket’s interior area |
MDP Modeling
Each subtask is modeled as a Markov Decision Process \((S, A, P, R, \gamma)\). The state space is defined by preprocessed visual observations from the robot’s head camera. Original RGB images of size \(160 \times 120\) are converted to grayscale, resized to \(84 \times 84\), and stacked over four consecutive frames to capture temporal dynamics. The stacked grayscale frames are then normalized to the range \([0,1]\), producing an input tensor of shape \(4 \times 84 \times 84\). The stacking of frames is essential because a single frame cannot fully reveal the motion state of the robot or the movement of the ball.
The action space for each subtask is discrete and consists of robot motions that can be executed by the Gait Manager or Motion Manager. For subtask 1, the actions are forward, turn left, and turn right. Subtask 2 includes forward, backward, turn left, turn right, left-arm catch, and right-arm sweep. Subtask 3 includes forward, turn left, turn right, and shoot. The motion parameters for the mobile actions are detailed in the following table:
| Action | setXAmplitude | setAAmplitude | setYAmplitude |
|---|---|---|---|
| Forward | 0.3 | 0 | 0.5 |
| Backward | -0.3 | 0 | 0.5 |
| Turn Left | 0.3 | 0.3 | 0.5 |
| Turn Right | 0.3 | -0.3 | 0.5 |
The reward function for each subtask follows a sparse reward design:
\[
R(s,a,s’) = \begin{cases} +100, & \text{if subtask goal is achieved}\\ -10, & \text{if maximum step limit is exceeded}\\ -5, & \text{if the ball is knocked down}\\ 0, & \text{otherwise} \end{cases}
\]
If the robot successfully completes the subtask, it receives a positive reward of 100. Exceeding the episode length of 30 steps results in a negative reward of -10. Knocking the ball off its pole gives a penalty of 5. All other actions receive zero reward.
Visual Representation Learning with CBAM
To improve the robot’s perception from raw pixels, I designed a lightweight visual representation module based on a convolutional neural network enhanced with the Convolutional Block Attention Module (CBAM). The CNN consists of three convolutional layers without pooling, preserving positional information that is crucial for perceiving object locations. The architecture is described in the following table:
| Layer | Input Channels | Output Channels | Kernel Size | Stride |
|---|---|---|---|---|
| Conv1 | 4 | 32 | 8 | 4 |
| Conv2 | 32 | 64 | 4 | 2 |
| Conv3 | 64 | 64 | 3 | 1 |
The output of the third convolutional layer is a feature map of size \(64 \times 7 \times 7\). This feature map is then passed through the CBAM module, which sequentially applies channel attention and spatial attention to emphasize task-relevant channels and spatial regions (e.g., the ball and basket) while suppressing background clutter. This design enhances the robustness and efficiency of perception without introducing excessive parameters.
PSR-SAC: Prioritized Successful Recent Experience Replay SAC
To improve sample efficiency and accelerate convergence, I developed a new experience replay mechanism called Priority Success Recent Experience Replay (PSR), which is integrated into the SAC framework. The intuition behind PSR is inspired by human learning: when acquiring a new skill, people benefit most from (a) successful attempts and (b) recent attempts. Based on this principle, the replay buffer system consists of three components:
- A global experience replay pool storing all transitions.
- A success experience replay pool storing transitions from successful episodes only.
- A temporary buffer storing the most recent episode’s transitions.
During each sampling step, a mini-batch is constructed by combining three types of data: samples from the recent episode buffer, prioritized samples from the global pool, and prioritized samples from the success pool. The priority of each sample is computed based on its TD-error:
\[
p_i = |\delta_i|^{\alpha} + \epsilon
\]
where \(\delta_i\) is the TD-error, \(\alpha\) controls the degree of prioritization, and \(\epsilon\) is a small constant ensuring non-zero priority. The sampling probability is:
\[
P(i) = \frac{p_i}{\sum_{k} p_k}
\]
To correct the bias introduced by prioritized sampling, importance sampling weights are applied:
\[
w_i = \left(\frac{1}{N \cdot P(i)}\right)^{\beta}
\]
where \(N\) is the experience pool size and \(\beta\) controls the correction strength.
For the discrete action space used in this chapter, the SAC update formulas are adapted accordingly. The soft Q-target is computed by summing over the full action distribution:
\[
\hat{Q}(s’, a’) = \sum_{a’} \pi_{\phi}(a’ | s’) \left[ Q_{\bar{\theta}}(s’, a’) – \alpha \log \pi_{\phi}(a’ | s’) \right]
\]
The Critic loss becomes:
\[
J_{Q}(\theta) = \mathbb{E}_{(s,a,r,s’)\sim\mathcal{D}}\left[\left(Q_{\theta}(s,a) – (r + \gamma \hat{Q}(s’,a’))\right)^2\right]
\]
And the Actor loss is:
\[
J_{\pi}(\phi) = \mathbb{E}_{s \sim \mathcal{D}}\left[\sum_{a} \pi_{\phi}(a|s)\left(\alpha \log \pi_{\phi}(a|s) – Q_{\theta}(s,a)\right)\right]
\]
The complete PSR-SAC procedure is summarized in the following algorithm block:
| Algorithm 1: PSR-SAC |
|---|
| Initialize Actor network \(\phi\), Critic networks \(\theta_1, \theta_2\), target networks \(\bar{\theta}_1, \bar{\theta}_2\) |
| Initialize global replay pool \(D_g\), success replay pool \(D_s\), recent episode buffer \(B_r\) |
| for \(t = 1, 2, \dots, T\) do |
| Select action \(a_t \sim \pi_{\phi}(\cdot | s_t)\) |
| Execute \(a_t\), obtain reward \(r_t\), next state \(s_{t+1}\), done flag \(d_t\) |
| Store \((s_t, a_t, r_t, s_{t+1}, d_t)\) in \(B_r\) |
| if done then |
| Move all data from \(B_r\) to \(D_g\) |
| if success then copy \(B_r\) data to \(D_s\) |
| Clear \(B_r\) |
| if training condition met then |
| Sample from \(B_r\), prioritize sample from \(D_s\), prioritize sample from \(D_g\) |
| Compute Critic loss, Actor loss, update networks |
| Soft update target networks |
Method 2: H-PPO for Humanoid Robot Basketball Shooting
Problem Analysis
The discrete action space approach in the previous chapter has a fundamental limitation when applied to the shooting subtask: the robot cannot adjust its shooting strength based on its distance to the basket. When the ball position is random, the robot-to-basket distance varies across episodes, and a fixed shooting strength inevitably leads to poor performance. Moreover, the sparse reward signal does not provide adequate guidance during the early phase of training, causing slow convergence.
Hybrid Action Space Modeling
To address this, I formulated the shooting subtask as a parameterized Markov Decision Process (PAMDP). In a PAMDP, each discrete action \(a_d \in \mathcal{A}_d\) is associated with a continuous parameter vector \(x \in \mathcal{X}_{a_d}\). A complete hybrid action is defined as the tuple \((a_d, x)\). For the basketball shooting task, the discrete action space is defined as: forward (0), turn left (1), turn right (2), and shoot (3). Each of these actions has an associated continuous parameter: for moving actions, the parameter controls execution duration in the range [0.2, 0.4] seconds; for the shooting action, the parameter controls the joint velocity in the range [10, 15], corresponding to the throwing strength. Since the ranges of these parameters differ significantly, I mapped the Actor network’s output from the standard range [-1, 1] to the actual parameter range using linear transformation:
\[
v = \frac{(u + 1)}{2} \cdot (v_{\max} – v_{\min}) + v_{\min}
\]
where \(u \in [-1,1]\) is the normalized output from the continuous Actor network, and [\(v_{\min}, v_{\max}\)] defines the physical range of the parameter.
Network Architecture for H-PPO
The H-PPO algorithm uses an Actor-Critic structure. The Critic network is a state-value network that outputs the value estimate for the current state. The Actor network is divided into two parallel branches:
- Discrete Actor network: outputs a probability distribution over the discrete actions, modeled as a categorical distribution.
- Continuous Actor network: outputs the mean and standard deviation of a Gaussian distribution over continuous parameters.
The overall architecture is described in the following table:
| Component | Structure | Output |
|---|---|---|
| Shared CNN | 3 convolutional layers (same as Section on CBAM without attention module) | 64 \(\times\) 7 \(\times\) 7 flattened feature |
| Critic Network | Full-connected layers with relu activation | State value \(V(s)\) |
| Discrete Actor | Full-connected layers with softmax activation | Action probabilities \(p(a_d | s)\) |
| Continuous Actor | Full-connected layers with tanh activation for mean, softplus for std | Parameter vector \(x\) |
The discrete action \(a_d\) is sampled from the categorical distribution, and the continuous parameter \(x\) is sampled from the Gaussian distribution. The executed hybrid action is \((a_d, x)\); for example, if the discrete action is “shoot” and the continuous parameter is 0.2, the actual joint velocity command is computed by inverse transformation.
Reward Design with Human Experience
To alleviate the sparse reward problem, I introduced a dense reward term based on human prior knowledge. From a human perspective, a basketball player is more likely to score when facing the basket. Therefore, I trained a YOLOv5n object detection model to detect the basket in the robot’s camera view during training. The base reward is augmented with a human experience reward that encourages the robot to orient itself toward the basket:
\[
R_{he} = \alpha_{he} \cdot \exp\left(-\frac{d_{center}^2}{2\sigma_{he}^2}\right)
\]
where \(d_{center}\) is the horizontal pixel distance between the center of the detected basket bounding box and the center of the image, \(\sigma_{he}\) controls the sharpness of the reward decay, and \(\alpha_{he}\) scales the reward magnitude. The total reward is:
\[
R_{total} = w_1 \cdot R_{base} + w_2 \cdot R_{he}
\]
where \(R_{base}\) is the sparse task reward defined earlier.
H-PPO Algorithm Design
The continuous or discrete PPO algorithms cannot directly handle hybrid action spaces. I therefore designed an H-PPO algorithm that extends PPO to parameterized action spaces. The GAE advantage estimator is used as follows:
\[
\hat{A}_t = \delta_t + (\gamma \lambda) \delta_{t+1} + \cdots + (\gamma \lambda)^{T-t} \delta_T
\]
The Critic network is updated by minimizing the value loss:
\[
J_{value} = \mathbb{E}_t\left[\left(V_{\psi}(s_t) – \hat{R}_t\right)^2\right]
\]
The discrete Actor network is updated with the clipped surrogate objective:
\[
J_{discrete} = \mathbb{E}_t\left[\min\left(r_t^d \hat{A}_t, \operatorname{clip}(r_t^d, 1-\epsilon, 1+\epsilon)\hat{A}_t\right)\right]
\]
where \(r_t^d = \frac{\pi_{\phi_d}(a_t^d | s_t)}{\pi_{\phi_d^{old}}(a_t^d | s_t)}\) is the importance sampling ratio for the discrete action. Similarly, the continuous Actor network is updated as:
\[
J_{continuous} = \mathbb{E}_t\left[\min\left(r_t^c \hat{A}_t, \operatorname{clip}(r_t^c, 1-\epsilon, 1+\epsilon)\hat{A}_t\right)\right]
\]
with \(r_t^c = \frac{\pi_{\phi_c}(x_t | s_t)}{\pi_{\phi_c^{old}}(x_t | s_t)}\).
Experimental Setup
All experiments were conducted in Webots R2021b using Python 3.7 and PyTorch 1.13.1. The hardware configuration includes an AMD Ryzen 9 7590X CPU, 32 GB RAM, and an NVIDIA GeForce RTX 3070 GPU running on Windows 10.
| Environmental Parameter | Value (unit: cm) |
|---|---|
| Robot start position | (90, 0) |
| Basket position | (0, 0) |
| Fixed ball position | (67, 0) |
| Random ball range (x-axis) | 60 to 70 |
| Random ball range (y-axis) | -23 to 23 |
| Basket height | 40 |
The training parameters for PSR-SAC are listed in the following table:
| Parameter | Value |
|---|---|
| Training steps | 30000 |
| Global replay capacity | 8192 |
| Success replay capacity | 4096 |
| Discount factor | 0.99 |
| Mini-batch size | 512 |
| Learning rate | 0.0003 |
| Priority exponent \(\alpha\) | 0.7 |
| Importance sampling weight \(\beta\) | 0.4 |
| Target network update frequency | 50 |
The training parameters for H-PPO are listed below:
| Parameter | Value |
|---|---|
| Training steps | 50000 |
| Simulation steps per policy update | 512 |
| Discount factor | 0.99 |
| Mini-batch size | 512 |
| GAE parameter | 0.95 |
| Epochs per update | 4 |
| Clipping coefficient | 0.1 |
| Entropy coefficient | 0.01 |
| Value loss coefficient | 0.5 |
| Gradient clipping max norm | 0.5 |
Experimental Results and Analysis
Overall Performance of PSR-SAC
To evaluate the overall effectiveness of the proposed CBAM-PSR-SAC algorithm, I compared it with three baseline algorithms: SAC, DQN, and PPO. Each algorithm was trained three times with different random seeds, and the average results are reported. The comparison was performed on three subtasks under both fixed and random ball position scenarios. The average episode reward curves show that CBAM-PSR-SAC achieves consistently higher rewards and faster convergence across all three subtasks.
The following table summarizes the number of episodes required for convergence across different algorithms:
| Algorithm | Fixed Subtask 1 | Fixed Subtask 2 | Fixed Subtask 3 | Random Subtask 1 | Random Subtask 2 | Random Subtask 3 |
|---|---|---|---|---|---|---|
| DQN | 600 | 670 | 520 | 1050 | 750 | 7800 |
| PPO | 780 | 1000 | 930 | 1330 | 1800 | 10000 |
| SAC | 400 | 120 | 300 | 1120 | 600 | 7000 |
| CBAM-PSR-SAC | 380 | 120 | 280 | 980 | 420 | 6500 |
In terms of convergence speed, the proposed CBAM-PSR-SAC algorithm achieves improvements over SAC of 5%, 6.7%, 12.5%, 30%, and 7.1% respectively for the five subtask-scenario combinations where SAC also converged. The success rates measured over 1000 test episodes are shown below:
| Algorithm | Fixed Subtask 1 | Fixed Subtask 2 | Fixed Subtask 3 | Random Subtask 1 | Random Subtask 2 | Random Subtask 3 |
|---|---|---|---|---|---|---|
| DQN | 95.1 | 56.6 | 96.5 | 90.1 | 56.8 | 55.6 |
| PPO | 82.5 | 92.5 | 98.6 | 82.9 | 48.5 | 41.3 |
| SAC | 92.7 | 99.6 | 98.2 | 93.4 | 60.1 | 48.8 |
| CBAM-PSR-SAC | 99.6 | 99.9 | 99.3 | 96.8 | 68.4 | 58.2 |
The improvement in success rates for the random ball scenario is particularly notable: an increase of 3.4 percentage points for subtask 1, 8.3 percentage points for subtask 2, and 9.4 percentage points for subtask 3 compared with SAC.
Ablation Study on Visual Representation Learning
To verify the effect of the CBAM attention module on representation learning, I compared CBAM-SAC with several other representation learning methods: VAE-SAC, DrQ-SAC, and ResNet-SAC, all evaluated on the random ball position scenario. The results show that CBAM-SAC converges faster and achieves higher average episode rewards in all three subtasks. The following table presents the success rates:
| Algorithm | Subtask 1 | Subtask 2 | Subtask 3 |
|---|---|---|---|
| SAC | 93.1 | 60.1 | 48.8 |
| DrQ-SAC | 91.2 | 47.5 | 42.1 |
| ResNet-SAC | 92.5 | 49.8 | 32.5 |
| VAE-SAC | 49.4 | 36.2 | 38.4 |
| CBAM-SAC | 96.3 | 65.8 | 56.2 |
These results demonstrate that the integration of the CBAM module improves both the convergence speed and the final task success rate for the humanoid robot basketball shooting task. The lightweight attention mechanism avoids the excessive parameter overhead introduced by deeper networks or complex generative models.
Ablation Study on Experience Replay Mechanism
To verify the effectiveness of the PSR sampling strategy, I compared PSR-SAC with PER-SAC, ER-MS-SAC, and PEC-SAC in the random ball position scenario. The PSR-SAC algorithm consistently achieves higher average episode rewards and faster convergence. The sensitivity to the success experience sampling ratio was also evaluated, with a value of 0.1 yielding the best trade-off.
The success rates for each experience replay variant are summarized below:
| Algorithm | Subtask 1 | Subtask 2 | Subtask 3 |
|---|---|---|---|
| SAC | 93.4 | 60.1 | 48.8 |
| PER-SAC | 95.8 | 66.2 | 52.8 |
| ER-MS-SAC | 94.2 | 53.6 | 49.6 |
| PEC-SAC | 93.3 | 45.2 | 46.5 |
| PSR-SAC | 96.4 | 67.5 | 57.9 |
The improvement of PSR-SAC over standard SAC is 3 percentage points for subtask 1, 7.4 percentage points for subtask 2, and 9.1 percentage points for subtask 3, validating the positive effect of combining prioritized sampling, success experience, and recent experiences.
Overall Performance of YOLO-H-PPO
In the shooting subtask, I compared the proposed YOLO-H-PPO with DQN, PPO, and CBAM-PSR-SAC. The convergence episodes and success rates are listed in the following tables. For the fixed ball position scenario, YOLO-H-PPO converges in 780 episodes compared to 930 episodes for PPO, a 16.1% improvement. For the random ball position scenario, YOLO-H-PPO converges in 5000 episodes, while DQN, PPO, and CBAM-PSR-SAC require 7800, 10000, and 6500 episodes, respectively.
| Algorithm | Fixed Ball Position | Random Ball Position |
|---|---|---|
| DQN | 520 | 7800 |
| PPO | 930 | 10000 |
| CBAM-PSR-SAC | 280 | 6500 |
| YOLO-H-PPO | 780 | 5000 |
The success rates and average number of steps per episode are as follows:
| Algorithm | Success Rate (Fixed) | Success Rate (Random) | Average Steps (Fixed) | Average Steps (Random) |
|---|---|---|---|---|
| DQN | 96.5 | 55.6 | 4.62 | 5.56 |
| PPO | 98.6 | 41.3 | 6.55 | 8.42 |
| CBAM-PSR-SAC | 99.3 | 58.2 | 3.95 | 6.27 |
| YOLO-H-PPO | 99.1 | 75.3 | 3.68 | 4.63 |
The YOLO-H-PPO algorithm improves the success rate by 17.1 percentage points over CBAM-PSR-SAC and by 34 percentage points over PPO in the random ball position scenario. Moreover, the robot is able to achieve the shooting goal in fewer action steps, indicating more efficient and targeted behavior.
Ablation on Hybrid Action Space Design
I compared H-PPO with MP-DQN (a hybrid action space algorithm) and discrete action space algorithms (DQN, PPO, CBAM-PSR-SAC) under the same sparse reward setting. The results show that H-PPO achieves a significantly higher success rate in the random ball position scenario compared to the discrete action space approaches. The advantage arises because H-PPO allows the robot to adjust both the discrete action selection and the associated continuous parameters (e.g., shooting strength), which is essential for dynamic task scenarios.
Ablation on Human-Experience Reward Design
To verify the contribution of the human-experience reward module, I compared PPO and H-PPO with and without this module in the random ball position scenario. The convergence episodes are summarized below:
| Algorithm | Sparse Reward | Human-Experience Reward |
|---|---|---|
| PPO | 6200 | 4000 |
| H-PPO | 8100 | 6700 |
The human-experience reward module improves convergence by 35.5% for PPO and by 17.3% for H-PPO. The success rates were largely maintained or slightly improved, as shown below:
| Algorithm | Sparse Reward | Human-Experience Reward |
|---|---|---|
| PPO | 41.3 | 43.2 |
| H-PPO | 73.6 | 75.3 |
From the qualitative analysis of the trained behavior, the robot at the early training stage struggles to orient itself correctly toward the basket and often shoots with inappropriate strength. After 10000 training episodes, the robot learned to adjust its body orientation and shooting strength according to its current position, successfully completing the shooting subtask in the random ball position scenario.
Discussion and Comparison of Key Findings
The experimental results reveal several important insights. First, the PSR-SAC method demonstrates that combining attention-based representation learning with a well-designed experience replay mechanism can substantially improve sample efficiency and convergence speed for humanoid robot basketball shooting tasks. The CBAM module helps the robot identify task-relevant features in the visual input, while the PSR strategy ensures that high-value experiences are sampled preferentially, thereby reducing redundant training and accelerating learning.
Second, the comparison between the discrete action space methods and the hybrid action space method shows that action space modeling plays a critical role in the success of complex manipulation tasks. In the random ball position scenario, the robot cannot rely on a fixed shooting strength because the distance to the basket changes. The hybrid action space design permits the robot to simultaneously choose the discrete action and the associated continuous parameter values, enabling more flexible and adaptive behavior. The improvement in success rate from 58.2% (CBAM-PSR-SAC) to 75.3% (YOLO-H-PPO) highlights the importance of this modeling choice.
Third, the human-experience reward design, implemented through the YOLOv5-based target detection module, provides a practical mechanism for injecting prior knowledge into the reinforcement learning pipeline. By rewarding the robot for keeping the basket near the center of its field of view, the exploration space is greatly reduced, allowing the algorithm to converge in significantly fewer episodes. This suggests that a well-designed dense reward function can be as important as the algorithm itself, especially in tasks where sparse rewards provide little guidance.
Finally, the integration of the object detection model into the reward design demonstrates a clear pathway for incorporating modern computer vision techniques into reinforcement learning-based robot control. This approach can be generalized to other tasks where spatial relationships between the robot and target objects are important.
Conclusion and Future Work
In this thesis, I investigated deep reinforcement learning methods for humanoid robot basketball shooting. I proposed two algorithm frameworks, PSR-SAC and H-PPO, to address the key challenges of sample efficiency, convergence speed, action space modeling, and sparse rewards. The PSR-SAC method incorporates the CBAM attention mechanism into visual representation learning and introduces a priority success recent experience replay mechanism to improve sample utilization. The H-PPO method introduces a hybrid action space design that enables the humanoid robot to simultaneously choose discrete actions and their continuous parameters, along with a human-experience-based reward function implemented using YOLOv5 target detection. The proposed methods enable the humanoid robot to learn basketball shooting skills through autonomous interaction with the simulated environment, demonstrating the potential of deep reinforcement learning for dynamic manipulation tasks.
The experimental results demonstrate that the proposed methods significantly outperform baseline algorithms in both convergence speed and task success rate. In the random ball position scenario, the PSR-SAC algorithm improved the shooting subtask success rate by 9.4 percentage points over SAC, and the YOLO-H-PPO algorithm further improved it by 17.1 percentage points over PSR-SAC. These results validate the effectiveness of the proposed framework for humanoid robot basketball shooting.
Looking forward, several directions warrant further investigation. First, the proposed methods have been validated only in simulation. Transferring the learned policies to a physical Robotis OP2 humanoid robot would require addressing the sim-to-real gap through techniques such as domain randomization and system identification. Second, the lower-level action primitives (e.g., walking and shooting motions) are still derived from hand-crafted controllers. Training these primitive skills directly using reinforcement learning could further improve the autonomy and generality of the humanoid robot. Third, the current hybrid action space decomposes the action into discrete and continuous parts but does not model correlations between different discrete actions and their parameters. A more advanced action representation that encodes discrete and continuous components jointly could yield more sample-efficient learning. Fourth, integrating more advanced perception models, such as semantic segmentation or depth estimation, may enhance the robot’s understanding of the scene and lead to more robust manipulation skills. Finally, extending the proposed methods to other dexterous manipulation tasks beyond basketball shooting, such as object grasping and assembly, would be a natural next step to test the generalizability of the framework.
In summary, this thesis contributes to the field of humanoid robot skill learning by proposing efficient deep reinforcement learning frameworks tailored to basketball shooting. The hybrid action space design and the human-experience reward shaping approach provide transferable paradigms for a broad range of humanoid robot manipulation tasks in dynamic environments.
