The ocean, encompassing vast resources and strategic importance, is progressively becoming a focal domain for global technological advancement. In alignment with strategic national directives to strengthen maritime capabilities, the independent development of underwater robotics has witnessed significant breakthroughs. Among these, bionic robots that mimic the morphology, locomotion, and behavioral characteristics of aquatic fauna stand out. These bionic platforms offer remarkable operational flexibility and exhibit minimal environmental disturbance, presenting not only superior underwater mobility but also promising application prospects across various fields.
In nature, fish schools demonstrate astonishing collective capabilities in activities such as foraging, predator evasion, and migratory navigation. Behaviors like orcas hunting seals or sailfish corralling sardines are quintessential examples of efficient decision-making and swarm intelligence. These biological phenomena have captivated researchers from diverse fields including control theory, bionics, and swarm intelligence, inspiring solutions for engineered systems. Particularly, the predator-prey inspired pursuit-evasion problem has emerged as a central research theme. For multi-robot systems, this task involves a team of pursuers coordinating to capture one or more evaders within a defined environment, such as underwater. This scenario represents a classic and challenging testbed for multi-agent cooperation, demanding real-time data processing, dynamic planning, and coordinated control.
Marine collective hunting behaviors provide a rich source of inspiration for developing cooperative strategies for multi-bionic robot systems. By emulating the interactive mechanisms observed in real fish schools, such artificial systems can achieve higher mission efficiency, showing immense potential in seabed exploration, security patrols, and ecological monitoring. Early work on robotic pursuit often relied on precise dynamical models or game-theoretic approaches, which struggle with nonlinearities and the high uncertainties inherent in complex, dynamic environments like the ocean. Furthermore, the computational cost of these traditional methods can grow prohibitively with the number of agents.
Reinforcement Learning (RL), and particularly Multi-Agent Reinforcement Learning (MARL), has risen as a powerful paradigm that foregoes the need for explicit, accurate models and learns effective policies through interaction. However, applying MARL to underwater bionic robots presents unique hurdles. These bionic robots are typically underactuated systems with complex hydrodynamics, and their motion is susceptible to disturbances from water currents. Developing efficient, reliable, and cooperative pursuit policies for a team of such bionic robots in a highly nonlinear, perturbed underwater environment remains a formidable challenge.
To address this, we propose a MARL-based training framework specifically designed for the collaborative pursuit task using underwater bionic robots. Our approach meticulously accounts for the unique locomotion characteristics of bionic robotic fish. We employ the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm, which operates under a centralized-training-with-decentralized-execution paradigm. This allows agents to leverage global information during training to learn better cooperative policies while acting based solely on local observations during execution, ensuring scalability and practicality. We design tailored state and action spaces that reflect the bionic robot’s motion constraints, craft a sophisticated reward function to incentivize effective teaming and pursuit, and implement a curriculum learning strategy to facilitate stable training convergence. Experimental validation in both simulation and real-world pool environments confirms the effectiveness and practicality of our proposed strategy.

Platform and Task Formulation for Aquatic Bionic Systems
The core of our experimental work is a custom-developed bionic robotic fish, whose design draws inspiration from the biomechanics of sharks. The bionic robot’s propulsion is generated by an oscillatory tail fin, mimicking thunniform swimming. This bio-inspired locomotion mechanism offers advantages in efficiency and maneuverability compared to traditional rotary thrusters, making it ideal for agile underwater operations.
The motion control of the bionic robot is governed by an artificially engineered Central Pattern Generator (CPG) model. The CPG produces smooth, periodic signals that drive the tail’s sinusoidal oscillation. By modulating the parameters of this CPG network, we can control the robot’s swimming speed and heading direction. The fundamental oscillation for each joint \( i \) can be modeled as a phase-coupled oscillator:
$$
\phi_i = 2\pi f t + \psi_i
$$
$$
r_i = A_i \sin(\phi_i) + O_i
$$
where \( f \) is the common frequency controlling speed, \( A_i \) is the amplitude, \( \psi_i \) is the phase offset determining the body wave, and \( O_i \) is the baseline offset crucial for steering. The turning radius \( R_{min} \) is directly influenced by the offset parameter \( O_i \) and is a critical constraint in motion planning. The key specifications of our bionic robotic fish are summarized in the table below.
| Attribute | Specification |
|---|---|
| Length | 0.68 m |
| Mass | 3.05 kg |
| Max Speed | 0.85 BL/s (Body Lengths per second) |
| Min Turning Radius (Pursuer) | 0.15 m |
| Min Turning Radius (Evader) | 0.12 m |
| Actuation | Multi-joint oscillatory tail |
| Control Core | STM32 Microcontroller |
We formulate the multi-bionic robot pursuit-evasion task within a bounded 2D aquatic plane. The pursuing team consists of \( N \) bionic robotic fish (in our experiments, \( N=2 \)), and the evading team has \( M \) agents (here, \( M=1 \)). To increase the challenge and reflect realistic asymmetries, the evader is granted superior agility, characterized by a smaller minimum turning radius and a higher maximum angular velocity. The primary objective for the pursuers is to collaboratively intercept the evader. A capture is deemed successful if any pursuer closes within a specified capture distance \( d_{capture} \) of the evader. The evader’s goal is to avoid this indefinitely or until a time limit expires.
The state transition for each bionic robot is managed via a waypoint tracking scheme. Given the periodic motion, the target waypoint for the next control step must be placed at a distance that accounts for the stride length per tail beat cycle. The kinematics between waypoints can be discretized. Let the current target pose be \( \mathbf{p}_t = (x_t, y_t, \theta_t) \) and the next be \( \mathbf{p}_{t+1} = (x_{t+1}, y_{t+1}, \theta_{t+1}) \). The required change in heading \( \Delta\theta \) and the step distance \( \Delta d \) form the basis of our action space. A backstepping controller is then used to enable the bionic robot to reliably track these successive waypoints, translating high-level MARL decisions into stable low-level locomotion.
A Multi-Agent Reinforcement Learning Framework for Bionic Teams
We adopt the MADDPG algorithm as the foundation of our learning framework. MADDPG is an actor-critic method adapted for multi-agent settings. Each bionic robot, or agent, has its own actor network (policy) \( \mu_i \) that maps its local observations \( o_i \) to a continuous action \( a_i \). Crucially, each agent also has a critic network \( Q_i \) that estimates the value of the joint action \( \mathbf{a} = (a_1, …, a_N) \) given the full state of the environment \( \mathbf{s} \). During training, the critics have access to this global state \( \mathbf{s} \) and all actions \( \mathbf{a} \), enabling them to learn a coordinated evaluation of the team’s behavior. The actors, however, are trained only with their local observations, ensuring that the resulting policy is executable in a decentralized manner. The core update rule for the critic of agent \( i \) is derived from the temporal-difference error:
$$
y = r_i + \gamma Q_i^{\mu’} (\mathbf{s}’, a_1′, …, a_N’) |_{a_j’=\mu_j'(o_j)}
$$
$$
\mathcal{L}(\theta_i^{Q}) = \mathbb{E}_{\mathbf{s}, \mathbf{a}, r, \mathbf{s’}}[(Q_i(\mathbf{s}, a_1, …, a_N) – y)^2]
$$
The actor policy for agent \( i \) is then updated by ascending the gradient of its expected return, approximated by the critic:
$$
\nabla_{\theta_i^{\mu}} J \approx \mathbb{E}_{\mathbf{s}, \mathbf{a}}[\nabla_{a_i} Q_i(\mathbf{s}, a_1, …, a_N) \nabla_{\theta_i^{\mu}} \mu_i(o_i)]
$$
This framework is ideal for our bionic robot pursuit task as it balances the need for coordinated learning with decentralized execution.
State Space Design for the Bionic Pursuit Task
The local observation \( o_i \) for each bionic robot agent must contain sufficient information for individual decision-making while fostering cooperation. For a pursuer \( i \), the state vector is carefully constructed:
$$
o_i = [x_i, y_i, \theta_i, v_i, \omega_i, \mathbf{d}_{i,*}, \mathbf{\phi}_{i,*}, \mathbf{d}_{i,ev}, \mathbf{\phi}_{i,ev}]
$$
This includes:
1. Ego State: Its own global position \( (x_i, y_i) \), heading \( \theta_i \), linear velocity \( v_i \), and angular velocity \( \omega_i \).
2. Team Awareness: Relative distances \( \mathbf{d}_{i,*} \) and bearings \( \mathbf{\phi}_{i,*} \) to all other pursuing bionic robots. This is crucial for maintaining formation and avoiding collisions.
3. Target Information: The relative distance \( \mathbf{d}_{i,ev} \) and bearing \( \mathbf{\phi}_{i,ev} \) to the evader. If the evader is outside the agent’s sensor range, these values are clipped or set to a default.
The evader’s observation space is similarly defined but includes relative information to all pursuers instead. This comprehensive observation allows each bionic robot to perceive the local geometric context of the pursuit, enabling the emergence of complex group strategies.
| State Component | Description | Dimension |
|---|---|---|
| Ego Pose & Velocity | \([x, y, \theta, v, \omega]\) | 5 |
| Teammate Rel. States | Distances \(d\) and bearings \(\phi\) to \(N-1\) teammates | \(2 \times (N-1)\) |
| Target Rel. State | Distance \(d_{ev}\) and bearing \(\phi_{ev}\) to evader | 2 |
| Total Dimension | For N=2 pursuers | 9 |
Action Space for Bionic Locomotion
The action output from the actor network must be directly mappable to the bionic robot’s control inputs. As discussed, high-level control is achieved by setting the next waypoint. Therefore, we define a 2-dimensional continuous action space for each bionic robot:
$$
a_i = [\Delta \theta_i, \Delta d_i]
$$
where:
– \( \Delta \theta_i \in [-\pi/4, \pi/4] \) rad is the immediate change in heading angle. This constraint reflects the mechanical limits of the bionic robot’s turning capability.
– \( \Delta d_i \in [\Delta d_{min}, \Delta d_{max}] \) meters is the step distance to the next waypoint. The minimum step \( \Delta d_{min} \) is set in relation to the agent’s minimum turning radius to ensure kinematic feasibility. Typically, \( \Delta d_{min}^{pursuer} > \Delta d_{min}^{evader} \), subtly reflecting the evader’s agility advantage.
| Agent Type | Action \(\Delta\theta\) | Action \(\Delta d_{min}\) | Rationale |
|---|---|---|---|
| Pursuer | \([-\pi/4, \pi/4]\) rad | 0.15 m | Allows sharp but feasible turns; step linked to larger turn radius. |
| Evader | \([-\pi/3, \pi/3]\) rad | 0.10 m | Larger heading change & smaller step enables tighter, more agile evasion. |
Crafting the Reward Function for Cooperative Pursuit
The design of the reward function \( r_i \) is pivotal for shaping the desired cooperative behavior among the bionic robots. We structure it as a weighted sum of several components:
$$
r_i^p(t) = w_c \cdot r_c + w_d \cdot r_d + w_b \cdot r_b + w_s \cdot r_s
$$
1. Capture Reward (\( r_c \)): A large positive reward is granted to every pursuer upon successful capture. This is the primary extrinsic goal.
$$ r_c = R_{success} \quad \text{if } \min_j ||\mathbf{p}_j – \mathbf{p}_{ev}|| < d_{capture} $$
2. Distance Reduction Reward (\( r_d \)): To encourage continuous pursuit, pursuers receive a reward proportional to the reduction in distance to the evader compared to the previous step.
$$ r_d = \Delta d_{t-1 \to t}^{i, ev} $$
3. Formation & Inter-agent Bonus (\( r_b \)): To promote intelligent cooperation, an additional bonus is given when pursuers position themselves strategically relative to both the evader and each other (e.g., maintaining a certain angular separation around the target).
$$ r_b = f(\phi_{i,ev}, \{\phi_{k,ev}\}) $$
4. State Penalties (\( r_s \)): Negative rewards (penalties) are imposed for undesirable states, such as colliding with another bionic robot or moving out of the predefined operational boundary.
$$ r_s = P_{collision} \quad \text{or} \quad P_{boundary} $$
The reward for the evader is simpler: a large positive reward for surviving until the episode time limit, a large negative reward for being captured, and a small penalty for approaching the boundary too closely to encourage central evasion.
| Reward Component | Agent | Typical Value | Purpose |
|---|---|---|---|
| Capture Success | Pursuer | +100.0 | Primary goal incentive |
| Distance Reduction | Pursuer | Scale: 0.1 per meter | Encourage closing in |
| Cooperative Bonus | Pursuer | +5.0 | Incentivize flanking/pincer moves |
| Collision Penalty | All | -20.0 | Ensure safe operation |
| Boundary Penalty | All | -2.0 per step | Keep agents in playable area |
| Survival Bonus | Evader | +50.0 (episodic) | Evader’s primary goal |
Curriculum Learning for Stable Training
Training a multi-bionic robot system from scratch in a challenging pursuit task can be unstable. The pursuers may fail to find the evader initially, receiving no meaningful reward signal. To mitigate this, we employ a simple yet effective curriculum learning strategy. Training starts with a “slow” evader, whose speed and agility parameters are reduced. This gives the pursuing bionic robots a higher chance of incidental success, allowing the policy to receive the capture reward and begin learning. Gradually, over many episodes, the evader’s capabilities are increased until they reach their full, superior level. This progressive increase in task difficulty guides the learning process and leads to more robust and effective final policies for the bionic robot team.
Experimental Validation and Strategy Analysis
We validated our MARL framework through extensive experiments in a simulated environment and with real bionic robotic fish in a 5m × 4m indoor pool. The real-world setup uses an overhead global vision system for positional feedback, which is sent to each bionic robot’s onboard controller to execute the waypoint-based actions derived from the trained policy.
The training process showed clear learning progression. Initially, the pursuing bionic robots moved randomly. As training advanced under the curriculum, they learned to chase the evader’s current position. Finally, the mature policy exhibited sophisticated, cooperative behaviors. The key hyperparameters for the MADDPG training are listed below.
| Hyperparameter | Value |
|---|---|
| Actor Learning Rate | 1e-4 |
| Critic Learning Rate | 1e-3 |
| Replay Buffer Size | 1e6 |
| Batch Size | 512 |
| Discount Factor (γ) | 0.95 |
| Soft Update Rate (τ) | 0.01 |
| Curriculum Stages | 5 |
The emergent strategies were compelling. The two pursuing bionic robots did not merely engage in independent direct chases, which often failed against the more agile evader. Instead, they demonstrated clear coordination:
1. Pincer Movement: The pursuers learned to approach the evader from divergent angles, effectively reducing its escape routes. One bionic robot would apply pressure from the front or side, while the other maneuvered to cut off the anticipated evasion path, leading to a successful intercept in a corner.
2. Role Specialization (Emergent): Although not explicitly programmed, one bionic robot often took a more aggressive “chaser” role, while the other adopted a “blocker” or “interceptor” role, positioning itself strategically based on the evader’s motion relative to the chaser.
Quantitative results over 100 evaluation episodes with the final policy against the full-capability evader are summarized below. We compare our MADDPG-based strategy against two baselines: a simple heuristic pursuit law (each pursuer heads directly toward the evader) and a pre-MARL RL method trained with Independent DDPG (where each agent learns with its own critic unaware of others’ actions).
| Metric / Strategy | Heuristic Pursuit | Independent DDPG | Our MADDPG |
|---|---|---|---|
| Capture Success Rate | 22% | 41% | 89% |
| Average Capture Time (s) | 45.2 | 32.7 | 18.3 |
| Team Collision Rate | 8% | 15% | 2% |
The results clearly demonstrate the superiority of our MARL framework. The high success rate and significantly reduced capture time underscore the effectiveness of the learned cooperative policy. The low collision rate further confirms that the bionic robots have learned to coordinate their movements intelligently, not just with the target but also with each other. The real-world pool experiments successfully replicated these cooperative behaviors, proving the practical deployability of the policy trained largely in simulation. The robustness of the bionic robot’s low-level tracking controller allowed the high-level MARL policy to transfer effectively.
Conclusion and Future Perspectives
In this work, we have presented a comprehensive MARL framework for enabling effective collaborative pursuit in a team of underwater bionic robots. By explicitly accounting for the unique locomotion characteristics of oscillatory-fin bionic robotic fish in the design of the state-action space and reward function, and by leveraging the centralized-training-decentralized-execution paradigm of MADDPG, we have successfully trained policies that exhibit sophisticated, emergent cooperative strategies. The policies go beyond simple chasing, demonstrating advanced concepts like flanking, interception, and role adaptation. Experimental validation in both simulation and real aquatic environments confirms the strategy’s high performance, robustness, and practicality.
This research provides a significant step forward in the intelligent control of multi-bionic robot systems for complex underwater tasks. The ability of these bio-inspired machines to learn and execute cooperative hunting strategies opens new avenues for applications in autonomous underwater monitoring, resource surveying, and environmental interaction where agile, coordinated group action is essential. The success of this bionic robot application also reinforces the value of MARL as a tool for solving complex, dynamic multi-agent problems in the physical world.
Future work will focus on several key extensions. First, we aim to enhance the perceptual autonomy of the bionic robot by integrating onboard sensors like vision or sonar, moving away from reliance on global positioning systems. Second, we plan to scale the system to larger teams of bionic robots (\( N > 2 \)) and multiple evaders, which will introduce new challenges in coordination and assignment. Third, investigating more complex, dynamic, and cluttered underwater environments with obstacles will test the robustness and generalization of the learned policies. Finally, continual learning and adaptation mechanisms will be explored to allow the bionic robot team to adjust its strategy online when faced with novel evader behaviors or changing environmental conditions.
