A newly released review in Computer Engineering and Applications examines how unmanned aerial vehicles can autonomously search for odor sources by combining airborne sensing, source-location cognition, search decision-making, and coordinated action. Authored by Jia Yingmiao, Fan Shurui, and Wang Li of Hebei University of Technology, the review positions the problem as a representative embodied intelligence challenge: an agent must perceive a physical environment, maintain an internal probabilistic understanding of a hidden source, choose actions under uncertainty, and revise its beliefs through movement and feedback.
The review argues that autonomous odor source search for UAVs is not merely a gas-detection task. It is an embodied intelligence process that links observation, cognition, decision-making, and collaboration into a closed loop. That framing places olfaction alongside vision, language, touch, and manipulation as a sensorimotor channel through which an embodied intelligence can understand and act in the physical world. The review notes that embodied intelligence research has expanded rapidly into multimodal perception, vision-language-action models, navigation, obstacle avoidance, grasping, and collaborative operations, yet olfactory information remains comparatively underexplored because it is sparse, intermittent, and strongly affected by wind fields and platform motion.

The review is organized around a four-part chain: observation, cognition, decision, and collaboration. It surveys embodied olfactory observation, embodied search cognition and decision-making, and collective embodied collaboration. It also discusses current research priorities and future directions for information-driven olfactory autonomous search. According to the review, probabilistic inference offers a clear way to connect odor observations, source-location beliefs, and search actions, because a posterior probability distribution can represent uncertainty about the source position while information gain, entropy reduction, divergence, or reward functions can guide action selection.
1. Embodied Intelligence Reframes UAV Olfactory Search
In the review, embodied intelligence refers to the capacity of a physical agent to perceive and interact with the real world, learn from that interaction, and make decisions accordingly. Under this definition, a UAV searching for an odor source must do more than detect a chemical. It must use onboard sensors to acquire local environmental information, combine that information with its own position and motion state, infer where the source may be, and select actions that improve its knowledge of the environment. This is a continuous loop in which sensing, cognition, decision-making, and execution influence one another.
The review emphasizes that odor source localization has traditionally been studied through gas sensing, mixed-gas identification, mobile robot source localization, or general cooperative mechanisms. However, those perspectives often treat sensing, estimation, and control as separate modules. An embodied intelligence perspective instead asks how local odor observations become decision-relevant information, how source-location beliefs are updated over time, how actions are evaluated under physical constraints, and how multiple UAVs share observations and coordinate behavior. The review therefore presents information-driven olfactory search as a way to organize the field around the closed-loop relationship between observation, cognition, decision, and action.
The authors note that information-driven methods based on probabilistic inference are especially suitable for this analysis. They can represent potential source locations through posterior probability distributions and choose search actions using measures such as information gain, entropy reduction, divergence, or reward. This makes it possible to describe how an embodied intelligence gradually reduces uncertainty while moving through a turbulent and partially observable environment.
2. Olfactory Search as a Defining Embodied Intelligence Problem
The review describes odor source search as a classic embodied intelligence problem because the agent cannot rely on a static map or a single instantaneous measurement. The chemical plume is often intermittent, the wind field changes, and the source may be hidden or inaccessible. The UAV must therefore move, sample, update its beliefs, and act again. The physical body of the UAV matters: rotor downwash, flight dynamics, sensor placement, vibration, electromagnetic interference, temperature, and humidity all affect what the agent can perceive and how it can move.
This perspective extends embodied intelligence beyond visual and linguistic tasks. Most embodied intelligence research has concentrated on explicit sensory channels such as vision, language, touch, and manipulation. Olfaction, by contrast, provides sparse and intermittent information that is easily disturbed by airflow and platform motion. Bringing olfaction into embodied intelligence can broaden an agent’s ability to perceive non-visual information and can offer new research opportunities for UAV autonomous search and odor source localization.
The review also observes that existing surveys have contributed valuable classifications of algorithms, system architectures, and robot collaboration. However, many remain focused on sensor systems, algorithm categories, or general cooperative mechanisms. They pay less attention to the closed-loop relationship among odor observation, source-location cognition, search decision, and cooperative action in UAV-specific settings. The review therefore aims to reorganize the field from the perspective of an embodied intelligence loop, using the information-driven paradigm as a central analytical thread.
3. Research Activity Signals a Field Moving Toward Intelligent Autonomy
The review reports a literature search in the Web of Science database using terms such as odor source localization, gas source location, autonomous search, drones, unmanned aerial vehicles, and embodied intelligence. The results indicate overall growth in related publications across the past fifteen years, though different research areas have developed at different speeds. Odor source localization research began earlier and has formed a relatively stable foundation. Autonomous search and UAV research accelerated after 2016, with UAV publications growing especially quickly. Embodied intelligence research started later, began a slow phase around 2019, and then showed a clear increase in 2024, reflecting growing attention to the field.
For the review, this trend matters because it suggests that olfactory autonomous search is moving toward intelligent autonomy. UAVs offer high mobility and strong environmental adaptability, making them attractive platforms for tasks that require sensing, movement, and decision-making in complex spaces. At the same time, embodied intelligence provides a conceptual framework for integrating perception, cognition, and action rather than treating them as isolated technical problems.
4. Embodied Olfactory Observation: The Information Entry Point
In the embodied intelligence loop, odor observation is the starting point. It is the information entry through which the UAV acquires environmental data and begins to form source-location cognition. During search, the UAV relies on an onboard artificial olfactory system to obtain odor concentration, wind direction, wind speed, and other environmental parameters. It combines those measurements with its own position and motion state, turning instantaneous sensor responses into evidence for source inference. The review identifies three main topics in this layer: airborne gas observation systems, sensor response processing, and odor observation representation.
4.1 Airborne Gas Observation Systems
An airborne gas observation system typically consists of a gas sampling module, gas sensors or a sensor array, positioning and communication modules, and a data acquisition and processing unit. The UAV uses this system to obtain gas concentration, sensor responses, or multidimensional odor features, while recording the spatial location where odor signals appear. In autonomous olfactory search, the purpose of observation is not only to determine whether a target gas is present. It is also to provide evidence for embodied decisions. Therefore, the performance of an onboard observation system cannot be judged only by sensitivity or accuracy; it must also be assessed by whether the observed information can support source inference and search decisions.
Gas sensors and sensor arrays provide the most basic observations. Common sensors include metal oxide sensors, electrochemical sensors, nondispersive infrared sensors, photoionization detectors, and particulate matter sensors. They are used for pollutant monitoring, leak detection, and gas concentration measurement. Sensor arrays or electronic nose systems can obtain multidimensional odor information through multiple sensitive units, which helps with gas identification and feature extraction in complex odor environments.
For autonomous search, decision-making usually requires more than gas concentration. It also needs synchronized spatial position, wind field, and environmental parameters. Some studies integrate sensor modules, ultrasonic anemometers, meteorological sensors, and positioning modules on the same UAV platform to collect concentration, wind field, weather, and spatial information simultaneously. Such multisource observation is closer to the needs of autonomous search because it links local odor responses with environmental information and strengthens the UAV’s ability to judge the value of future search regions.
The review also highlights rotor downwash as a major challenge for multirotor UAVs. The airflow generated by rotors changes the gas movement around the aircraft, creating a discrepancy between measured concentration and true environmental concentration. Existing solutions fall into two categories. One approach adjusts sensor or sampling inlet positions to keep sampling points away from the rotor disturbance zone, for example through side-extending placement, high top-mounted inlets, long suspended sampling tubes, or lowered independent measurement platforms. The other approach uses active sampling structures, such as micro pumps, inlet pipes, and sampling chambers, to draw target gas into a sensing unit and stabilize detection. Additional measures include independent power supply, metal shielding, constant-temperature packaging, drying tubes, and temperature compensation to reduce electromagnetic interference and environmental effects.
4.2 Sensor Response Processing and Concentration Estimation
Raw sensor outputs cannot directly serve search decisions. Gas sensors usually produce voltage, current, resistance changes, or multichannel responses that are affected by noise, response lag, environmental variation, and platform motion. The review explains that sensor responses must be processed into concentration observations, concentration trends, or stable odor features that can be associated with spatial position, wind field information, and search actions. This transformation allows raw odor responses to support plume detection, source-location belief updating, and search direction selection.
In controlled environments or scenarios with a known target gas, sensor response and gas concentration can be mapped through calibration curves or empirical models. Common methods include linear regression, polynomial fitting, exponential models, power-law models, and sensitivity-based response functions. These methods are simple and computationally efficient, making them suitable for real-time operation on UAVs. However, fixed calibration relationships are difficult to maintain over time because temperature, humidity, and sensor drift affect measurements. Periodic calibration is therefore necessary, and concentration estimation for search tasks cannot rely only on static calibration.
Electronic nose systems obtain multidimensional odor responses through multiple sensitive units, providing richer gas features than a single sensor. Traditional pattern recognition methods can process electronic nose signals. These methods often use principal component analysis, linear discriminant analysis, or other feature extraction techniques to reduce raw sensor data, then combine K-nearest neighbors, support vector machines, random forests, or support vector regression to map response features to gas categories or concentrations. Some studies also introduce swarm intelligence optimization for feature selection or model parameter tuning to improve concentration prediction in mixed-gas conditions.
When odor environments become more complex, neural network methods are increasingly used to model nonlinear mappings from sensor responses to concentration observations. Convolutional neural networks can extract spatial correlation features from multisensor response matrices or odor images. One-dimensional convolutional neural networks are suitable for time series. Long short-term memory networks and gated recurrent units can capture dynamic sensor responses over time. Hybrid models such as CNN-LSTM and CNN-GRU combine spatial feature extraction with temporal dependency modeling and have shown adaptability in mixed-gas concentration prediction, early identification, and continuous response modeling.
The review stresses that concentration estimation is not the final goal. It is an input to autonomous search decisions. Different search methods use concentration information in different ways, so concentration must be converted into usable decision information according to the requirements of each method.
4.3 Odor Observation Representations
Different autonomous search methods use odor information differently. Odor observation can be represented as concentration intensity, concentration gradient, hit events, or continuous response sequences. The representation determines how much odor information is preserved and influences source-location belief updating and action selection.
Hit events are a common representation in information-driven odor source search. Because turbulent plumes are intermittent and sparse, sensors often receive valid odor signals only at certain positions or times. Information-driven algorithms therefore often abstract gas detection into discrete hit events. The most common approach sets a threshold based on sensor response or concentration: a value above the threshold is recorded as a hit, and a value below it is recorded as a miss, forming a binary observation. Some studies count hits over a sampling period and use hit frequency to describe the strength of odor contact, providing richer temporal information than binary observations.
Hit-event representations are often used in Infotaxis and related information-driven search methods. They support source-location posterior updating, are computationally simple, and tolerate some sensor fluctuation and calibration error. However, hit events discard concentration amplitude, response rate, and continuous temporal features. When the plume is highly nonuniform or sensor response lags, relying only on hit events may underuse available observation information.
Concentration intensity is the most direct representation. It is usually obtained from sensor responses through calibration, compensation, or concentration estimation models. Concentration gradient represents the spatial or temporal trend of concentration. Both are continuous odor observations and preserve more amplitude information than hit events. Search methods based on concentration intensity or gradient are essentially local optimization processes. The UAV judges the direction of increasing or decreasing odor and tends to move toward regions of higher concentration. These representations are widely used in gradient search, chemotaxis, reactive plume tracking, and some heuristic search methods, especially when odor distribution is relatively continuous, the wind field is stable, or the search scale is small.
Their advantages are intuitive form and simple computation. They can quickly convert odor observations into local motion directions and are useful during plume discovery and short-term tracking. However, in open environments, odor plumes are affected by turbulent diffusion, wind fluctuations, and UAV motion. A local concentration peak does not necessarily point to the true source. Concentration intensity and gradient are therefore better suited as local motion adjustment information. To support global source judgment, they need to be combined with wind field, spatial position, and historical observations.
With the development of sensor arrays and learning models, odor observation can also be represented as continuous sensor responses or multichannel time series. Compared with instantaneous concentration readings, time-series responses preserve gas exposure, response rise, recovery, and drift. In probabilistic inference methods, continuous concentration observations or sensor responses can be used to construct observation likelihoods, allowing source-location posterior updating to use observation amplitude rather than only detection events. For learning-driven search methods, continuous responses can be combined with wind speed, wind direction, position, and attitude to form observation vectors for policy learning and action selection. One study used a gas sensor array and anemometer time-series data with a CNN-LSTM model to estimate the distance between the sensor and the source, showing that continuous sensor array data can further support source-localization decisions.
Compared with hit events and concentration gradients, time-series responses preserve more complete odor dynamics and are more suitable for continuous observation modeling, learning-driven search, and multimodal information fusion. However, they are higher-dimensional and place greater demands on real-time inference, training sample coverage, and cross-environment generalization. For UAV odor source search, the choice of observation representation depends not only on sensor measurement capability but also on source-location cognition, action decision-making, and multi-UAV collaborative information fusion.
5. Embodied Search Cognition and Decision-Making
In embodied olfactory search, cognition and decision-making connect odor observation with physical action. The UAV must form a probabilistic understanding of potential source locations from local, sparse, and uncertain odor observations, then convert that cognitive state into executable search actions. The review classifies existing methods into gradient-based and reactive search methods, swarm intelligence optimization methods, probabilistic inference-based information-driven methods, and learning-driven methods. The review focuses on information-driven olfactory autonomous search, while treating the other categories as important foundations for understanding the field.
Gradient-based and reactive methods reflect early local perception-action coupling in olfactory search. Swarm intelligence optimization methods reflect parallel exploration and heuristic collaboration among multiple agents. Learning-driven methods represent a more recent direction in which search strategies are learned from data and interaction experience. The review then examines information-driven methods, especially Infotaxis-like approaches, to analyze how source-location cognition is represented, updated, and transformed into embodied search decisions.
5.1 Other Typical Olfactory Autonomous Search Methods
The review summarizes gradient-based and reactive search methods according to how actions are generated. These include reactive search based on typical behavioral patterns, chemotaxis based on concentration gradient and wind direction, and gradient search based on extremum-seeking control. From an embodied intelligence perspective, these methods emphasize immediate responses to local environmental stimuli. Their advantages are simple structure, strong real-time performance, and ease of implementation on resource-constrained platforms. Their limitation is that they mainly rely on local instantaneous observations and lack continuous modeling of historical odor information and source-location uncertainty. As a result, their adaptability is limited in intermittent plumes, complex wind fields, and three-dimensional UAV search scenarios.
Swarm intelligence optimization methods treat odor source search as a spatial global optimization problem. The core idea is to construct a fitness function from odor observations and use multiple search individuals to update positions and exchange information until the optimal location is found. Typical methods include particle swarm optimization, whale optimization algorithm, grey wolf optimizer, and wind-direction artificial ecosystem-based optimization. From an embodied intelligence perspective, these methods extend local search by a single agent to parallel exploration by multiple embodied individuals, reflecting distributed cooperative search. Their strength is strong global search ability, which is suitable for multi-platform cooperative detection. However, performance depends heavily on the fitness function and parameter settings. They also lack explicit representation of dynamic odor fields, multimodal environments, and source-location uncertainty, while complex hybrid mechanisms can increase online implementation difficulty.
Learning-driven odor source search methods use reinforcement learning, deep neural networks, or multimodal learning models to learn mappings from observation states to search actions through environmental interaction or training data. Typical methods include deep Q-networks, particle clustering deep Q-networks, dueling deep Q-networks, reinforcement learning fuzzy systems, gated recurrent unit proximal policy optimization, convolutional neural networks, long short-term memory networks, vision-olfaction fusion networks, odor compasses, and large language model-assisted decision-making. From an embodied intelligence perspective, learning-driven methods emphasize that an agent acquires search capability through perception-action interaction. They can automatically extract state features from odor responses, historical trajectories, wind direction, obstacles, and visual information, and generate search strategies. Their advantage is strong potential adaptability to complex environments and multimodal observations. However, they often depend on large training datasets and high-quality simulation environments. Policy performance is strongly affected by state representation, reward functions, and training scenarios, and generalization, interpretability, and safety in real UAV platforms still require validation.
| Method category | Basic idea | Typical methods | Main advantages | Limitations |
|---|---|---|---|---|
| Reactive search based on typical behavioral patterns | Execute preset behaviors according to odor presence or concentration changes | E. coli, spiral, hex-path, adaptive bio-inspired navigation | Simple structure and strong real-time performance | Search performance depends on empirical parameters |
| Chemotaxis based on concentration gradient and wind direction | Determine movement direction from local concentration differences and wind direction | Concentration gradient search, bilateral sensor chemotaxis, upwind search, concentration-wind fusion search | Intuitive decision-making and strong interpretability | Depends on the stability of local concentration gradient and wind information |
| Gradient search based on extremum-seeking control | Approach concentration extremum through perturbation, feedback, and gradient estimation | GA-ESC, PSO-ESC | Better stability than direct gradient search | Mainly depends on local structure of the concentration field |
| Method category | Basic idea | Typical methods | Main advantages | Limitations |
|---|---|---|---|---|
| Typical swarm intelligence optimization algorithms | Find the source region through fitness evaluation and swarm position updates | PSO, WOA, GWO, WAEO | Strong global search ability and suitability for parallel search | Depends on fitness function and parameters, prone to premature convergence, limited adaptability to dynamic odor fields and multimodal environments |
| Hybrid swarm intelligence algorithms | Enhance environmental adaptability by integrating obstacle avoidance, cooperation, or adaptive mechanisms | PSO-APF, adaptive-weight PSO, swarm intelligence-emotion learning | Balances search efficiency, flight safety, and complex environment adaptability | Complex structure, many parameters, and high online implementation difficulty |
| Method category | Basic idea | Typical methods | Main advantages | Limitations |
|---|---|---|---|---|
| Reinforcement learning methods | Learn mappings from observation states to search actions through environmental interaction and optimize long-term cumulative rewards | DQN, PC-DQN, dueling DQN, reinforcement learning fuzzy systems | Reduces dependence on manually designed motion rules and suits discrete action decisions, continuous control, and partially observable search tasks | Depends on extensive interaction training; policy performance is strongly affected by state representation, reward functions, and training environments |
| Deep neural network methods | Use neural networks to extract state features from odor responses, historical observations, and environmental information and support action generation or direction estimation | GRU-PPO, CNN, LSTM, vision-olfaction fusion networks, odor compass, LLM-assisted decision-making | Strong feature learning ability and can fuse spatial features, time series, and multimodal observations | Depends on sufficient training data and high-quality simulation environments; generalization, interpretability, and safety in real UAV platforms still require validation |
5.2 Probabilistic Inference-Based Information-Driven Search
Probabilistic inference-based information-driven search formulates odor source localization as active perception and decision-making under uncertainty. The method continuously updates a posterior probability distribution over source locations using historical observations, then selects the next movement direction according to how much a candidate action is expected to reduce source-location uncertainty. Because it explicitly represents source-location uncertainty and actively selects high-information actions, this approach is well suited to sparse observations, intermittent plumes, and large-scale turbulent environments.
From the perspective of the embodied intelligence loop, information-driven search clearly connects odor observation, source-location cognition, search decision, and action. In 2007, Vergassola and colleagues proposed Infotaxis, which replaced concentration with information as the decision basis. Instead of directly tracking instantaneous concentration gradients, Infotaxis updates a posterior probability distribution over source locations and selects actions that most reduce uncertainty. Moraud and colleagues later applied Infotaxis to a robot platform and validated its effectiveness and robustness under sparse odor conditions. Zhang and colleagues further analyzed Infotaxis performance in sparse environments, discussing how search distance and path-unit form affect search performance. This showed that spatial discretization and path-unit design influence the practical behavior of information-driven strategies.
The review describes Infotaxis as a relatively complete observation-cognition-decision-action framework. The agent obtains local odor observations, combines them with wind field and position information, and updates a posterior probability map. It then predicts possible observation outcomes for different candidate actions and evaluates their effect on the posterior distribution. Finally, it selects the action with the greatest expected information gain, gradually improving its knowledge of the source location through movement.
5.3 Source-Location Cognition Representation and Updating
Source-location cognition representation and updating are key to information-driven search. They convert local odor observations obtained during movement into a probabilistic expression of source location. In classical Infotaxis, the source-location posterior is often represented as a grid-based probability map. This divides the search area into finite cells, which facilitates numerical computation and action evaluation, but computational complexity grows quickly with spatial resolution and search area. In continuous spaces, three-dimensional environments, and multiparameter estimation tasks, grid-based methods can face the curse of dimensionality.
To address the limitations of discrete grid representation in continuous spaces, Barbieri and colleagues extended Infotaxis to continuous space by describing source-location uncertainty through continuous probability density and information potential fields. Ristic and colleagues further expanded the unknown state from a single source position to joint estimation of source position and source strength, using sequential Monte Carlo methods to approximate and recursively update the posterior distribution. Compared with grid-based probability maps, sequential Monte Carlo methods can represent potential source terms through a limited number of weighted particles, offering greater flexibility in continuous state spaces and multiparameter estimation tasks. Particle filters and related methods have therefore become important tools for posterior cognition modeling in information-driven search.
As real robot platforms and complex sensor observations were introduced, posterior updating also moved from idealized probabilistic recursion toward integration with actual observation mechanisms. Hutchinson and colleagues built an observation model using continuous outputs and intermittent response characteristics of metal oxide gas sensors on a low-cost mobile robot, enabling online source-term estimation and information-driven search from real sensor observations. Chen and colleagues extended classical Infotaxis to particulate matter source localization, constructing probability maps for different particle-size modes from PMS7003 sensor data and proposing a multiprobability map fusion strategy to improve source localization under multimodal particulate observations. Related research has also combined particle filters with bio-inspired anemotaxis behavior, showing that posterior updating is no longer only a mathematical recursion over source-location probability but is increasingly combined with sensor response characteristics, observation modalities, and motion behaviors.
On this basis, source-location cognitive states have expanded from a single source-location probability map to richer environmental representations. Wang and colleagues constructed a plume distribution map on top of a source-location probability map, combined with historical wind field information to describe possible plume propagation regions, and achieved joint representation of source-location cognition and plume spatial distribution cognition. Heinonen and colleagues modeled Bayesian olfactory search as a partially observable Markov decision process and formalized the source-location posterior distribution as a belief state, providing a stricter theoretical framework for policy optimization in information-driven search. Sun and colleagues approached the problem from generative modeling, building a Gaussian mixture plume model to describe the time-varying transport process of continuous releases and combining posterior probability maximization with model residual minimization for source-location inference.
From the embodied intelligence loop perspective, source-location cognition representation and updating correspond to the transformation from odor observation to source-location cognition. The agent continuously revises its cognitive state based on local observations, providing the basis for information gain calculation, reward function design, and search action selection. This moves olfactory autonomous search from simple local reactions to active decision-making based on cognitive states.
5.4 Action Evaluation for Embodied Action
After source-location cognition representation and updating, the agent must convert its current cognitive state into executable search actions. This is the transition from cognition to action in the embodied intelligence loop. In classical Infotaxis, candidate actions are evaluated by expected information gain, that is, by predicting the posterior entropy change that different actions may cause and selecting the direction that most reduces source-location uncertainty. In practice, however, when several candidate actions cause similar entropy changes, the value differences among actions decrease. Near the source or when local probability distribution fluctuates strongly, relying only on entropy reduction may cause unstable movement direction. Later research therefore shifted from a single entropy-reduction criterion to richer action evaluation mechanisms, allowing search decisions to connect source-location cognition with embodied action more stably.
To address the limitations of a single Shannon entropy reduction criterion, studies have introduced different information measures to redefine action reward. Ristic and colleagues compared original Infotaxis, Infotaxis II, and reward functions based on Bhattacharyya distance under a sequential Monte Carlo framework, showing that information gain is not necessarily limited to posterior entropy reduction. Song and colleagues introduced local probability reliability to reduce the influence of unreliable probability regions on search direction when local probability perturbations affect action judgment. He and colleagues used Rényi divergence to describe the difference between prior and posterior distributions, enhancing the distinguishability of information among candidate actions. Hutchinson and colleagues proposed Entrotaxis, which changes action evaluation from minimizing expected posterior entropy to maximizing the entropy of the predicted observation distribution, simplifying decision computation while improving search efficiency. Zhao and colleagues further proposed RE-Entrotaxis, using historical observations to regress source strength and updating only source location through particle filtering, thereby reducing the computational burden of joint estimation. These methods focus on improving the expression ability and action discrimination ability of information reward measures so that candidate action evaluation is no longer limited to a single entropy-reduction form.
As search scenarios expanded from discrete planes to continuous spaces and obstacle environments, pursuing information gain alone could lead the agent toward high-uncertainty regions. Action evaluation therefore expanded from pure information reward calculation to joint consideration of candidate action generation, path cost, and spatial reachability. Park and colleagues proposed GMM-Infotaxis, using a Gaussian mixture model to generate candidate sampling actions with greater information value in continuous space and improving search performance on UAV platforms. An and colleagues incorporated path cost into the entropy-reduction reward and proposed RRT-Infotaxis, enabling information-driven search in complex obstacle environments to consider both uncertainty reduction and motion reachability. Loisy and colleagues proposed Space-Aware Infotaxis, introducing the average distance from the robot to the source into the reward function so that action selection balances information acquisition and source-region approach. Jiu and colleagues combined a source likelihood map with an improved artificial potential field to calculate AUV heading and used a re-discovery strategy after plume loss, reflecting a trend toward integrating information evaluation with motion constraints.
As search scenarios become more complex, search decisions must also regulate exploration and exploitation across different search stages. In the early stage, the agent needs to expand observation coverage to reduce global uncertainty. As the posterior distribution converges, decisions should focus more on fine approach to high-probability regions. Liu and colleagues proposed ASAInfotaxis II, using swarm intelligence optimization to adaptively adjust movement range and weight parameters, achieving dynamic balance between exploration and exploitation. Song and colleagues constructed an information-driven search method based on minimum free energy, adjusting search behavior according to posterior distribution convergence. Guo and colleagues proposed Clutaxis, constructing high-density local belief regions and combining a spiral search pattern to coordinate global search and local approach, reducing decision complexity while improving search efficiency.
Action evaluation for embodied action transforms source-location cognition into physical action. Different information measures improve the discriminability of candidate actions’ contributions to cognitive updating. Path cost, spatial reachability, and artificial potential fields make action evaluation more consistent with real platform motion constraints. Exploration-exploitation adjustment enhances the search process’s adaptability to different cognitive stages. However, as reward terms increase, parameter selection and objective trade-offs become more complex. Information-driven search action evaluation has therefore evolved from early single entropy-reduction maximization to a comprehensive decision mechanism for real embodied tasks. It must balance computation, motion cost, localization accuracy, and search efficiency so that the agent can continuously and stably perform source localization in complex physical environments.
5.5 Embodied Search Expansion of Information-Driven Methods
The review traces how information-driven methods have moved from idealized models to real embodied tasks. Early information-driven search methods were often studied in open, obstacle-free, and regularized search spaces, with candidate actions given by fixed neighborhoods or discrete direction sets. In real odor source search, however, environments often contain walls, equipment, no-fly zones, or random obstacles, and gas diffusion is affected by spatial structure, ventilation conditions, and unsteady flow fields. These structures directly limit the agent’s feasible action set and may cause the agent to linger in local regions, reducing search efficiency and success rate.
To address this, many studies incorporate obstacle constraints into information-driven decision-making. The agent no longer selects actions only by information gain but performs posterior updating, action evaluation, and path selection in traversable regions, reducing ineffective movement and local wandering caused by obstacles. Zhao and colleagues constructed a structured environment model based on grid blocked cells and extended Infotaxis and Entrotaxis to IWFA and EWFA, enabling information-driven strategies to perform action constraints and path selection in environments with forbidden areas. Zhu and colleagues introduced a repeated-search penalty into a tracking strategy based on particle filtering and information entropy, reducing inefficient back-and-forth movement near obstacles and improving search robustness in complex environments.
As research progressed, path planning was directly integrated into search decision-making. An and colleagues combined rapidly exploring random trees with information-driven search, generating candidate path branches that satisfy multistep lookahead requirements in continuous space, allowing the agent to simultaneously avoid obstacles, collect information, and estimate source terms in complex urban obstacle environments. Luong and colleagues proposed a search framework combining environment decomposition and algorithm switching. Local regions use Infotaxis for information-driven search, while cross-region transitions use Dijkstra path planning, alleviating local wandering and cross-region search difficulties in complex structured environments. Kim and colleagues proposed a gas source localization framework combining an indoor Gaussian diffusion model with dual-mode information-driven planning. The local search stage uses RRT-Infotaxis for fine information collection, while deadlock or locally constrained scenarios switch to global planning based on frontier exploration. The introduction of path planning enhances global reachability and complex environment adaptability, allowing the agent to coordinate local information collection with cross-region movement. However, such methods usually require additional environment maps or spatial decomposition information and increase the computational cost of path generation, algorithm switching, and global planning. Planning results may also be affected by environment modeling errors.
Beyond obstacle constraints, information-driven search has also expanded toward hazardous environments, unsteady diffusion environments, and specific mission scenarios. For example, in hazardous gas leaks, Zhang and colleagues proposed the SSCEM framework for UAVs, which performs particle-filter source-term estimation, designs a continuous-domain heuristic action set, and constructs a cost function that considers exploitation, exploration, and cumulative exposure, balancing search efficiency, flight safety, and risk avoidance. For unsteady buoyant plume environments, Song and colleagues proposed an improved information-taxis source-tracing method based on exponential weight adjustment and stratified resampling, correcting weight normalization and combining residual information resampling to improve particle-set representation of dynamic posterior distributions. In nuclear accident radioactive leakage scenarios, Chen and colleagues introduced radioactive decay terms and cleaning factors into the hit-rate function of classical Infotaxis, allowing posterior estimation to adapt to dilution and decay in radiation environments. These studies show that information-driven search has expanded from ideal odor plume search to application scenarios with greater safety risks and physical complexity.
As search platforms expanded from ground robots to UAVs and underwater robots with three-dimensional motion capability, information-driven search also moved from two-dimensional planes to three-dimensional physical space. Odor plume diffusion in three-dimensional physical space is more spatially nonuniform. The agent must not only determine horizontal movement direction but also consider altitude changes and higher-dimensional posterior updating. Therefore, the key to three-dimensional information-driven search is not simply extending action directions from two dimensions to three dimensions, but rethinking three-dimensional plume modeling, candidate action generation, step-size adjustment, and stopping criteria.
Eggels and colleagues first tested and adapted Infotaxis in a turbulent three-dimensional channel flow environment, validating the feasibility of information-driven search in three-dimensional flow fields. Ruddick and colleagues further introduced three-dimensional plume models, three-dimensional candidate action sets, adaptive step-size strategies, and stopping conditions, and systematically validated three-dimensional Infotaxis through high-fidelity simulation and wind tunnel experiments, moving three-dimensional information-driven search from numerical adaptation toward experimental validation. Zhang and colleagues constructed a dynamic patrol information map for three-dimensional mountain environments, combining a fire existence probability map with an environmental uncertainty map to achieve three-dimensional UAV patrol planning under complex terrain, providing a reference for extending information-driven ideas to large-scale three-dimensional tasks. Three-dimensional expansion allows information-driven search to better adapt to real spatial tasks, while also bringing challenges such as expanded action spaces, increased posterior computational complexity, and difficult three-dimensional observation modeling.
Information-driven search expansion into complex environments and three-dimensional space significantly enhances its applicability to real mission scenarios, but also places higher demands on environment modeling accuracy, posterior updating efficiency, action planning capability, and online decision stability. In large-scale, highly uncertain, and sparsely observed search scenarios, a single agent is easily limited by finite sensing coverage and insufficient search efficiency. Research has therefore further extended information-driven search to multi-agent cooperative detection, hoping to improve source-location cognition efficiency and overall search performance through multi-point observation, information sharing, and cooperative decision-making.
| Technical stage | Improvement direction | Representative methods | Advantages | Limitations |
|---|---|---|---|---|
| Posterior cognition representation | Continuous-space representation | Continuous-space Infotaxis | Reduces dependence on discrete grids and improves adaptability to continuous space | Numerical computation and candidate action evaluation become more complex |
| Posterior cognition representation | Particle-filter representation | Particle-filter source-term estimation | Suitable for continuous state spaces and joint source-location/source-strength estimation | Particle degeneracy, resampling, and computational overhead |
| Posterior cognition representation | Composite cognitive representation | Plume distribution map, belief state, Gaussian mixture plume model | Expands source-location cognition expression and enhances description of plume propagation and dynamic environments | Requires more environmental information and historical wind field data |
| Posterior probability updating | Updating with real sensor observations | MOX sensor information-driven search, multimodal probability map method for particulate matter sensors | Makes posterior updating closer to real sensor responses | Depends on sensor calibration and observation model accuracy |
| Posterior probability updating | Particle-filter recursive updating | Sequential Monte Carlo source-term estimation | Suitable for nonlinear, non-Gaussian, and high-dimensional source-term estimation | Large computation; particle number affects estimation accuracy |
| Posterior probability updating | Combination with anemotaxis behavior | Particle-filter bio-inspired anemotaxis method | Combines probabilistic cognition with wind cues and improves plume tracking and source-region approach stability | Performance affected by wind observation accuracy and behavioral rules |
| Embodied action evaluation | Modifying information measures | Infotaxis II, Bhattacharyya distance, Rényi-Infotaxis, Entrotaxis, RE-Entrotaxis | Improves candidate action value discrimination and avoids the limitations of a single entropy-reduction criterion | Information measure parameter selection affects search performance |
| Embodied action evaluation | Fusing information measures with other factors | GMM-Infotaxis, RRT-Infotaxis, Space-Aware Infotaxis | Balances information acquisition and motion cost in action evaluation | Complex multi-objective trade-offs and difficult parameter setting |
| Embodied action evaluation | Adaptive exploration-exploitation adjustment | ASAInfotaxis II, minimum free energy, Clutaxis | Dynamically adjusts global exploration and local approach according to search stage | Adjustment rules are somewhat empirical and generalization requires validation |
| Embodied search expansion | Treating constrained regions as unreachable | IWFA, EWFA | Prevents the searcher from entering obstacles or impassable regions | Requires known environmental constraint information and is limited in flexibility |
| Embodied search expansion | Integrating path planning | RRT-Infotaxis, Dijkstra-Infotaxis | Improves reachability and global search ability in complex structured environments | Path planning and algorithm switching increase computational cost |
| Embodied search expansion | Expanding to special scenarios | SSCEM, Infotaxis under unsteady environments, Infotaxis for nuclear accident radioactive leakage | Adapts information-driven search to safety risks and special physical processes | Cost functions and task constraints become more complex |
| Embodied search expansion | Three-dimensional space search | Three-dimensional Infotaxis and action-space expansion | Adapts to three-dimensional search tasks for UAVs and underwater robots | Action space expands, and three-dimensional posterior updating and observation modeling become more difficult |
6. Collective Embodied Collaboration
When multiple UAVs participate in olfactory autonomous search, the embodied intelligence process expands from a single-agent loop to a collective loop of observation sharing, cognition fusion, and cooperative decision-making. The key is no longer simply parallel search by multiple platforms, but how to transform distributed local observations into consistent or complementary collective cognition and then support cooperative decisions.
The review notes that swarm intelligence optimization algorithms naturally involve parallel exploration and information exchange among multiple search individuals. When used for multi-robot or multi-UAV source search, they can form a multi-agent cooperative search process. Learning-driven methods can also learn cooperative behavior among UAVs through multi-agent reinforcement learning, multi-policy cooperation, or centralized training with decentralized execution. However, their core issues are mainly state representation, reward design, sample efficiency, policy generalization, and simulation-to-reality transfer. Their cooperative mechanisms are more often external behaviors of trained policies. The review therefore focuses mainly on information-driven cooperative source search methods.
6.1 Task Characteristics of Information-Driven Multi-UAV Cooperative Search
Information-driven UAV source search collaboration follows the chain of observation sharing, cognition fusion, and cooperative decision-making. Observation sharing is the foundation. Multiple UAVs sample in different spatial positions and exchange positions, trajectories, odor observations, wind speed, and wind direction, expanding plume detection coverage and improving collective environmental awareness. Cognition fusion is the core. Multiple UAVs can jointly construct source-location probability cognition from shared observations or fuse their local probability maps and particle weights to obtain a collective source-location judgment. Cooperative decision-making further acts on search behavior by sharing candidate actions, target regions, or task allocation results so that different UAVs form complementary searches, reduce repeated search and ineffective aggregation, and achieve cooperation between global exploration and local localization.
Multi-UAV cooperation is not simply about increasing the number of search agents. Without effective information sharing and task coordination, multiple UAVs may overlap search regions, produce redundant observations, aggregate prematurely, or interfere with one another. Multi-UAV systems also must handle flight safety, communication constraints, and real-time decision-making.
6.2 Observation Sharing
Observation sharing is the most basic form of cooperation in information-driven multi-agent source search. It incorporates local sensor information obtained by multiple agents during search into a shared source-location probability map update process, allowing the system to eliminate low-probability regions faster and improve reliability in judging potential source regions. In information-driven cooperative source search, observation sharing is usually the first step from single-agent posterior updating to multi-agent collective cognition.
Early multi-agent information-driven search mainly shared odor detection events to update a common posterior probability map. In 2009, Masson and colleagues extended Infotaxis to collective search, where multiple agents quickly exchanged odor detection results obtained within short time intervals and jointly constructed a shared posterior probability map. Each agent did not maintain an independent source-location cognition but treated its own observations as part of shared collective information. Zhang and colleagues further considered odor source search under limited spatial perception and proposed a multi-robot odor source search algorithm for weakly sensing robots. The method did not require a complete environment model or precise spatial perception and instead achieved cooperation by sharing path information, including positions where odor particles were detected and not detected, so that multi-point path observations jointly participated in source-location probability estimation.
On the basis of shared probability maps, Huang and colleagues argued that shared observations should not be weighted equally and that observation reliability should adjust their influence on collective probability cognition. The study used relative distance between agents and the pollution source and the number of sampling cues to construct confidence factors, correcting the shared probability map update process and improving probability map convergence speed and source-region judgment reliability.
As source-term estimation requirements increased, observation sharing expanded from odor hit events and path cues to multi-platform measurement sequences and their corresponding positions. Ristic and colleagues proposed a multi-robot information-driven autonomous search method for hazardous source search in turbulent environments. The method assumes fully connected communication among robots, allowing each platform to obtain all robots’ measurements and positions and perform replicated centralized Bayesian fusion based on the same global observation dataset. The study used multi-platform, multi-time sensor measurement sequences for posterior estimation of source parameters such as source location and release strength, making it more suitable for joint source-location and source-strength estimation. Park and Oh further linked observation sharing with distributed cooperation levels, proposing noncooperative, passive cooperative, and negotiated cooperative methods for multi-mobile-sensor source search and source-term estimation. Measurements obtained by multiple mobile sensors at the same time and different spatial positions were organized into joint observation vectors, and under the assumption of independent sensor noise, sensor observation likelihoods were multiplied to update potential source-term weights in particle filtering. Observation sharing improved source-term estimation accuracy and provided a foundation for higher-level cooperation based on action exchange and negotiation.
However, observation sharing alone cannot fully solve all problems in multi-agent cooperative source search. Hajieghrary and colleagues argued that simple observation sharing is not always sufficiently reliable. After multiple agents share detection results, uncertainty about source-location estimation can decrease rapidly, and the team may become overconfident in an incorrect source-location estimate or even make a false source-location judgment. Therefore, on the basis of observation sharing, it is necessary to further consider differences among different agents’ local cognition and how to fuse them.
6.3 Distributed Cognition Fusion
Distributed cognition fusion mainly integrates local source-location cognition formed by multiple agents. Each agent usually constructs a source-location probability map, particle weight distribution, or source-term parameter estimation based on its own observation history. Through posterior distribution fusion, particle fusion, cognition-difference weighting, consistency updating, or parameter exchange, distributed local cognition is integrated into a collective source-location judgment. The fusion object is usually not a single sensor observation but posterior probabilities, local probability maps, or source-term estimation results accumulated from observations.
Different agents may form different source-location cognitions because of their positions, observation histories, and odor detection results. Measuring and using cognitive differences among agents is therefore a key issue in distributed cognition fusion. Hajieghrary and colleagues proposed a cooperative search strategy based on KL divergence, comparing differences between source-location probability distributions of different agents to guide search, thereby reducing the risk of incorrect convergence from simple observation sharing while retaining the advantages of multi-point information acquisition. This idea extended KL divergence from information measurement to multi-agent cognitive cooperation and provided a foundation for subsequent fusion methods based on cognitive differences.
Socialtaxis is one representative method of distributed cognition fusion. In 2017, Karpas and colleagues proposed Socialtaxis, combining individual information gain in single-agent Infotaxis with collective social information. When making decisions, agents consider not only the reduction of uncertainty in their own posterior probability distribution but also differences between their distribution and those of other agents. Unlike methods that share one probability map, Socialtaxis does not require all agents to form completely consistent source-location cognition. Instead, it encourages individuals to use their own local posterior while referencing collective cognitive differences, maintaining better spatial dispersion and information diversity in the early search stage. Later research noted that Socialtaxis can achieve collective information-greedy search with low computational complexity and few information exchanges, but its original form also has problems such as bias toward exploration, insufficient cooperation, and lack of communication connectivity constraints.
To balance individual cognition and collective cognition, Song and colleagues proposed a multi-robot cooperative Infotaxis method based on cognitive differences. The method uses relative entropy to measure differences between different agents’ estimates of source-location distribution and assigns adaptive weights to sampling cues from different agents. During Bayesian updating, agents no longer indiscriminately accept all local observations but adjust the relative influence of their own cues and collective cues according to cognitive differences, allowing each agent to form a private source-location probability distribution consistent with its own observation history. Compared with simply sharing one probability map, this method preserves the independence of individual cognition and reduces the impact of a single agent’s false detection or local misjudgment on the entire collective posterior.
However, posterior fusion based on cognitive differences usually requires communication of probability maps, which can bring high computational and communication overhead in real-time search. To improve practical feasibility, related research has attempted to reduce probability map complexity using particle filtering or Gaussian fitting. Song and colleagues addressed the communication and computational burden of particle filtering in cooperative information-taxis methods and proposed a cooperative search method based on particle filtering and Gaussian fitting. Mean, covariance, and other Gaussian parameters approximate the full particle set, so agents only need to transmit key parameters to complete cognitive difference calculation and source-location cognition updating. Ling and colleagues also introduced cognitive-difference weighted measurement fusion and particle fusion in a multi-agent radioactive source search task, allowing observation information from different agents to selectively enter the fusion process according to cognitive correlation, and combined an adaptive step-size free-energy search strategy to improve search efficiency. This shows that distributed cognition fusion does not necessarily depend on full probability map exchange and can also be realized through particles, weights, or Gaussian parameters.
Most of these methods still require a certain degree of information exchange and posterior interaction. When the number of agents increases or communication conditions are limited, maintaining collective cognitive consistency in a decentralized structure becomes another key issue. Rahbar and colleagues proposed a distributed source-term estimation algorithm that enables a multi-robot system to update source-term cognition through a distributed posterior estimation mechanism without a centralized node. Ristic and colleagues further addressed decentralized multi-platform hazardous source search, proposing that each platform independently run a Rao-Blackwellised particle filter to achieve sequential source parameter estimation and use neighbor measurement exchange and consistency cooperative control to maintain formation and communication connectivity constraints during search. Nanavati and colleagues combined distributed Bayesian filtering with coverage control and proposed a consensus-based belief update mechanism in which each robot exchanges local posterior distributions with neighbors and fuses them using KL averaging, gradually forming a consistent source-term estimation belief across the network. Such methods do not require every platform to know all observations but instead form distributed estimates of source parameters through local communication and neighbor interaction, making them more suitable for cooperative search under communication constraints and scale expansion.
Distributed cognition fusion addresses how multiple agents integrate local source-location judgments. The above methods use posterior distribution differences, cognitive-difference weighting, particle fusion, Gaussian parameter exchange, and neighbor consistency updating to allow the collective to form more reliable source-location cognition while preserving individual observation differences. However, cognition-level fusion cannot directly guarantee complementary search behavior. If each agent still selects actions independently, repeated search or local aggregation may still occur. Therefore, action-level mechanisms such as candidate action exchange, target allocation, or negotiation are also needed.
6.4 Collaborative Decision-Making
After collective cognition is formed, if each agent still independently selects its next movement direction according to its own action evaluation function, the collective may still experience repeated search, local aggregation, insufficient coverage, or action conflicts. Collaborative decision-making therefore needs to interact around candidate movement directions, target regions, action values, or task allocation results so that multiple agents form more complementary search behavior at the action level.
The most direct collaborative decision-making approach is to construct a joint action space and uniformly evaluate candidate action combinations for all agents. Early collective Infotaxis already reflected this idea. Masson and colleagues, while proposing a shared posterior probability map, further discussed two cooperation levels: one in which agents share a source-location probability map but independently evaluate their own actions, and another in which the joint actions of the entire collective are evaluated so that agents achieve full cooperation during action selection. However, the computational cost of a collective joint action space grows exponentially with the number of agents and the number of candidate actions per agent, making full enumeration difficult to apply in real time in multi-agent scenarios.
To address the excessive computation of joint action spaces, Park and colleagues proposed a negotiated cooperation method. Instead of a central node enumerating all joint actions, each agent first calculates a locally optimal action based on its own source-term estimation and then exchanges candidate control decisions with other agents before execution. After obtaining other members’ temporary decisions, each agent assumes that other agents will move according to their current decisions and updates its estimate of future observations and information gain, then reevaluates its own action. This process iterates until decision changes stabilize or a maximum number of iterations is reached. Negotiated cooperation uses a local decision-decision exchange-decision update process to make individual actions gradually approach collective consistency or collective effectiveness. This method reduces computational complexity while retaining action interaction among different agents.
However, achieving action consistency through negotiation does not necessarily guarantee a reasonable spatial distribution. If multiple agents make decisions based on similar source-location cognition, even after negotiation they may overconcentrate in the same high-probability region, leading to insufficient spatial coverage. To address this, Luo and colleagues introduced a collaborative decision correction mechanism on the basis of cognitive-difference-driven source-term estimation. Based on shared source-location cognition, the method combines the deviation angle and distance of agents relative to the estimated source location to correct candidate actions under the Infotaxis II framework, allowing different robots to form a more reasonable spatial distribution around the potential source region. This method adds collective spatial relationships into action evaluation, enabling search behavior to consider both information acquisition and spatial dispersion and reducing ineffective aggregation.
When the posterior probability distribution contains multiple high-probability regions or multiple potential source hypotheses, collaborative decision-making must further address the question of who searches for which odor source. Bourne and colleagues proposed a Bayesian-bio-inspired fusion search method that provides another approach. The method feeds the posterior probability distribution back to bio-inspired search behaviors such as biased random walk and surge-casting, and uses a task allocation mechanism to guide different robots to verify different source-location hypotheses. In this way, multiple robots do not have to track the same highest-probability region simultaneously but can form search divisions according to different candidate source regions in the posterior distribution. This type of method emphasizes target region allocation and hypothesis verification and is suitable for scenarios with multimodal posterior distributions, high source-location uncertainty, or easy repeated search by multiple robots.
Collaborative decision-making focuses on how multiple agents can form more reasonable choices at the action level. Joint action evaluation can select a better action combination from the collective level. Negotiation iteration reduces the computational burden of centralized optimization through local decision exchange. Action correction helps maintain reasonable spatial dispersion among agents. Task allocation can guide different individuals to verify different source-location hypotheses and reduce repeated search.
| Cooperation mechanism | Basic idea | Representative methods and mechanisms | Main advantages | Limitations |
|---|---|---|---|---|
| Observation sharing | Share local odor observations, paths, and position information of multiple UAVs for source-location probability updating | Collective Infotaxis, multi-weak-perception robot odor source search, joint observation vector construction, observation confidence factor | Expands sensing coverage | Simple sharing may amplify false observations and make the collective overconfident in an incorrect source location |
| Distributed cognition fusion | Fuse local posteriors, particle weights, source-term parameters, or probability maps formed by different UAVs from historical observations | Socialtaxis, cognitive-difference weighting, Gaussian parameter exchange, consistency updating | Preserves individual observation differences while forming more reliable collective source-location cognition | Requires probability distribution comparison and information exchange, potentially increasing computation and communication overhead |
| Collaborative decision-making | Coordinate candidate actions, target regions, or task allocation based on collective cognition | Joint action evaluation, negotiation iteration, action correction, target region allocation | Reduces repeated search and local aggregation and improves spatial coverage and search efficiency | Joint action space is computationally large, and negotiation and task allocation mechanisms are complex to design |
7. Discussion and Outlook
The review draws several conclusions from the embodied intelligence perspective. First, odor observation is the foundation for an embodied intelligence system to form source-location cognition. Its value lies not only in obtaining gas concentration or identifying gas categories but also in providing environmental information that can be used for source-location estimation and search decisions. Odor observation for information-driven search should therefore be designed around the needs of cognition and decision-making, moving toward observation modeling oriented to search decisions.
Second, information-driven olfactory autonomous search is moving from idealized information-gain maximization toward cognitive decision-making that incorporates embodied constraints. In both posterior updating and action evaluation, algorithmic improvements increasingly consider practical executability, introducing obstacles, motion cost, and risk avoidance into the information-driven framework. This unifies cognitive effectiveness and action safety under embodied constraints. Future work needs to further study information-driven decision models for complex wind fields, dynamic obstacles, and real flight dynamics constraints.
Third, the core problem of multi-UAV cooperative search is not sharing more information but determining which information is worth sharing, when to share it, and how to share it under communication, motion, and task constraints. In multi-UAV cooperative source search, information sharing does not necessarily improve performance. When the number of UAVs increases or the sharing mechanism is poorly designed, local false detections, sensor noise, and wind disturbances may spread rapidly and be amplified, causing the collective to become overconfident in an incorrect source region or even converge incorrectly. Frequent exchange of observation data, posterior probabilities, or candidate actions also increases communication burden. In unstable links, limited bandwidth, or large-scale search scenarios, excessive communication may weaken real-time performance. Future research therefore needs to develop selective communication, event-triggered communication, and lightweight cognitive representation methods so that multiple UAVs can maintain reliable collective source-location cognition and complementary search behavior while reducing communication redundancy.
Fourth, the review focuses on information-driven autonomous search algorithms, but this does not mean such methods have become the absolute mainstream in olfactory autonomous search or that they are advantageous in all mission scenarios. The focus is chosen because information-driven methods clearly connect odor observation, source-location cognition, and search decision-making, making them suitable for analyzing UAV olfactory autonomous search from the embodied intelligence loop perspective. In fact, reinforcement learning, deep learning, multimodal fusion, and large-model-assisted decision-making have also shown strong development potential and can learn search strategies in complex environments from data and interaction experience. Future research should further promote the integration of information-driven methods with learning-driven methods, swarm intelligence methods, and multimodal perception models to improve the interpretability, adaptability, and practical deployment capability of UAV olfactory autonomous search systems.
Information-driven olfactory autonomous search provides an important theoretical foundation for UAV autonomous cognition and action decisions in unknown odor environments. Future work in this direction needs to further break through limitations such as idealized observation models, simplified motion assumptions, and heavy reliance on high-frequency communication. It should more tightly integrate odor sensing, source-location cognition, embodied constraints, cooperative communication, and learning adaptation mechanisms, promoting UAV olfactory autonomous search from information-theoretic optimal decision-making toward embodied intelligence search that is executable, interpretable, and cooperative in real environments.
8. Conclusion
The review systematically examines odor observation, single-agent information-driven autonomous search, and multi-UAV cooperative search from the perspective of embodied intelligence. It first summarizes how odor information is transformed from raw sensor input into effective observations that serve source-location cognition and search decisions. It then briefly reviews different types of olfactory autonomous search methods and focuses on the development of information-driven methods, discussing how they achieve autonomous search decisions through source-location probability representation, posterior cognitive updating, and action evaluation. Finally, it reviews multi-UAV cooperative information-driven search methods around observation sharing, distributed cognition fusion, and collaborative decision-making, analyzing how multi-UAV systems form collective source-location cognition and complementary search behavior under communication and motion constraints.
Overall, UAV information-driven olfactory autonomous search is a closed-loop process of continuous interaction among odor observation, source-location cognition, search decision-making, and embodied action. Future research needs to further address real UAV platforms, strengthen effective observation modeling, multimodal information fusion, information-driven decision-making under embodied constraints, and multi-UAV cooperation mechanisms under communication constraints. Such progress can move UAV olfactory autonomous search methods toward application in complex real-world environments and further establish embodied intelligence as a guiding framework for autonomous olfactory search.
