Embodied Intelligence Powers Multi-Robot Interactive Service System for Smart Cultural Tourism

A research team from China Mobile Communications Group Tianjin Co., Ltd. has proposed a multi-robot interactive service system for smart cultural tourism that places embodied intelligence at the center of service design, field deployment, and human-robot interaction. The system, developed by Dang Yuewen, Zhao Dongming, Lu Miao, You Xia, Li Xiaoheng, and Chen Jingyan, integrates a large language model, multimodal perception, cross-form robot collaboration, and a cloud-based cultural tourism knowledge service platform. It is designed to address persistent limitations in complex open tourism environments, including insufficient environmental perception, limited multimodal interaction, and weak collaborative service capability across different robot forms.

The proposed system combines humanoid robots, quadruped robots, wheeled robots, and a cloud knowledge service platform into a coordinated architecture. It supports cultural tourism knowledge question answering, visitor interaction, intelligent service decision-making, and robot-assisted cultural performances such as traditional Chinese group dance and Tai Chi demonstrations. The research team deployed and validated the system at the Tianjin Five Great Avenues Begonia Festival, where it performed robotic tour guidance, intelligent question answering, and group performance functions.

The application results indicate that the embodied intelligence-based multi-robot system can improve interaction experience and service capability in smart cultural tourism scenarios. According to the research, the system provides a practical reference for the application of embodied intelligence robots in cultural tourism and related public service fields.

1. Background: Embodied Intelligence and the Evolution of Tourism Robots

Artificial intelligence technologies are accelerating the transformation of cultural tourism scenarios toward intelligent and interactive service modes. In recent years, large language models, multimodal perception, and embodied intelligence technologies have continued to develop, pushing robot systems from traditional single-function equipment toward intelligent agents capable of environmental understanding, autonomous decision-making, and natural interaction. At the same time, the demand for smart cultural tourism construction has continued to rise. Scenic area guidance, visitor services, and cultural interaction scenarios have placed higher requirements on intelligent service capability.

Service robots, digital humans, and intelligent interaction systems have gradually been applied to exhibition hall guidance, scenic area interpretation, and large cultural tourism events. These applications provide a new technical path for upgrading smart cultural tourism services. However, many existing systems still rely on single-robot or fixed-process interaction models. In complex open scenarios, such systems often face limited environmental perception, insufficient multimodal interaction, and weak cross-form robot collaborative service capability.

The rise of embodied intelligence offers a broader framework for addressing these challenges. Unlike disembodied artificial intelligence, embodied intelligence emphasizes the coupling of perception, cognition, decision-making, and action in physical agents. In cultural tourism, this means robots must not only answer questions but also understand visitors, interpret scenes, move through crowds, coordinate with other robots, and deliver culturally meaningful performances. The research team argues that this shift is essential for moving from isolated robotic functions to integrated service ecosystems.

2. Related Technological Foundations

2.1. Vision-Language-Action Models and Embodied Intelligence

In the field of robot intelligence research, Vision-Language-Action technology, commonly known as VLA, provides a new path for autonomous robot interaction and complex task execution. The RT-2 model unifies visual language understanding and robot action control, enabling the transfer of internet knowledge to robot control capability. The PaLM-E model further integrates vision, language, and robot state information, improving multimodal perception and embodied reasoning in robots. These developments have helped robot systems acquire environmental perception, semantic understanding, and autonomous interaction capabilities in open scenarios, laying a technical foundation for intelligent service applications in smart cultural tourism.

Within this context, embodied intelligence serves as an organizing concept. It connects model-level advances in vision, language, and action with physical robot deployment. For smart cultural tourism, embodied intelligence is not limited to navigation or manipulation. It also includes the ability to interpret visitor intent, coordinate across robot forms, and adapt service behavior according to real-time scene conditions.

2.2. Existing Gaps in Cultural Tourism Robot Services

Although the smart cultural tourism field has already carried out research and applications involving intelligent guidance, digital human interpretation, and voice question answering, most existing systems still use single-robot or fixed-flow interaction modes. The research identifies several gaps:

  • Environmental perception capability is limited in complex open tourism scenes.
  • Multimodal interaction remains insufficient for natural visitor engagement.
  • Cross-form robot collaborative service capability is weak.
  • Service processes are often fixed rather than dynamically adapted.
  • Cultural interaction formats are frequently single and limited in appeal.

These limitations reduce the ability of robot systems to serve large crowds, respond to changing visitor needs, and cover multiple service points within a scenic area. The proposed system aims to address these issues through an embodied intelligence architecture that combines large language model reasoning, multimodal perception, and cross-form robot collaboration.

3. System Architecture: A Three-Layer Embodied Intelligence Framework

To overcome the limitations of traditional cultural tourism robot systems, the research team constructed an embodied intelligence-based multi-robot interactive service architecture for smart cultural tourism. The overall system adopts a three-layer design consisting of an interaction layer, an intelligent service layer, and a robot execution layer. It also uses a cloud-edge-end collaborative deployment model to enable multi-robot coordinated service in smart cultural tourism scenarios.

3.1. Interaction Layer: Visitor Access and Natural Engagement

The interaction layer is mainly responsible for visitor access and interaction services. It includes voice interaction, mini-programs, and visitor terminals. These modules support visitor demand input, information display, and real-time interaction. Visitors can interact naturally with the robot system through voice, text, and actions. This layer provides the first point of contact between the public and the embodied intelligence service system.

By supporting multiple access channels, the interaction layer helps the system serve different visitor groups and usage preferences. Some visitors may prefer spoken questions, while others may use text input or mobile terminals. The system is designed to accept these inputs and route them into the intelligent service layer for understanding and response generation.

3.2. Intelligent Service Layer: Large Language Models, Knowledge, and Multimodal Perception

The intelligent service layer consists mainly of a large language model, a cultural tourism knowledge base, multimodal perception modules, and task decision modules. The large language model is responsible for visitor intent understanding, cultural tourism knowledge question answering, and service content generation. The cultural tourism knowledge base stores scenic area information, historical culture, activity content, route recommendations, and public service information. The multimodal perception module fuses voice, vision, and environmental state information to achieve visitor behavior understanding and real-time interaction state perception.

In this architecture, embodied intelligence emerges from the interaction between model reasoning and physical context. The large language model does not operate in isolation. It is grounded by the knowledge base, informed by multimodal perception, and connected to robot actions through the execution layer. This grounding is a central design principle of the proposed embodied intelligence system.

3.3. Robot Execution Layer: Humanoid, Quadruped, and Wheeled Robots

The execution layer consists of humanoid robots, quadruped robots, and wheeled robots. Different robot forms differ in movement ability, environmental adaptability, and interaction capability. The system combines a multi-robot collaborative service mechanism to enable different robots to coordinate in tour guidance, interactive parades, and cultural performance scenarios.

Compared with traditional single-robot service models, the multi-robot collaborative architecture can give full play to the advantages of different robots in interaction, mobility, and scene adaptation. Humanoid robots are primarily responsible for tour guidance, visitor interaction, and cultural performance tasks. Quadruped robots are responsible for interactive parades and dynamic displays. Wheeled robots are responsible for mobile guidance and scene linkage services. Together, these robot forms extend the reach and flexibility of the embodied intelligence service system.

3.4. Cloud-Edge-End Collaborative Deployment

The system adopts a cloud-edge-end collaborative deployment model. The cloud side is mainly responsible for large model inference and cultural tourism knowledge services. The edge side handles real-time interaction processing and task scheduling. The robot end is responsible for environmental perception and interaction execution. This division allows the system to balance computational intensity, response latency, and local autonomy.

For embodied intelligence, cloud-edge-end collaboration is important because robots operate in dynamic public environments. They must respond quickly to visitors while also accessing rich knowledge and model capabilities. The architecture is designed to support this balance by distributing intelligence across cloud, edge, and robot endpoints.

4. Multi-Robot Collaborative Service Workflow

To realize intelligent interaction and collaborative service in complex cultural tourism scenarios, the system constructs a multi-robot collaborative service workflow based on perception, understanding, decision-making, interaction, and feedback. During service, robots first obtain visitor needs and scene state information through voice, vision, and environmental perception modules. The system then combines the large language model and cultural tourism knowledge base to complete visitor intent recognition, knowledge retrieval, and service content generation.

According to different service needs, the system dynamically schedules different forms of robots to complete tour guidance, visitor interaction, interactive parades, and robot performances. During robot execution, the system collects visitor interaction states and environmental feedback information in real time. It uses the multimodal perception module to dynamically update service status. When visitor needs or scene states change, the system can dynamically adjust interaction content and robot service strategies, forming a closed loop of real-time collaborative service in complex open scenarios.

4.1. Perception and Understanding in Open Tourism Environments

Perception is the foundation of embodied intelligence in the proposed system. Robots must recognize not only physical obstacles but also social and service-related cues. Voice input reveals visitor questions and intentions. Visual input can help identify visitor states and environmental conditions. Environmental state information provides context for task execution. The system fuses these inputs to build a richer understanding of the service situation.

The large language model then interprets visitor intent and generates service content. The cultural tourism knowledge base grounds this generation in scenic area information, historical culture, activity content, route recommendations, and public service data. This combination allows the system to support open-scene natural language question answering and multi-turn interaction.

4.2. Decision-Making and Dynamic Task Allocation

Decision-making in the system is not limited to selecting an answer. It also involves deciding which robot should perform which task, where, and in what sequence. The system considers task requirements, robot operating states, and scene information to achieve dynamic task allocation and collaborative service among different robot forms.

For example, a visitor may request route guidance, cultural interpretation, or an interactive performance. The system can assign a humanoid robot for explanation and interaction, a wheeled robot for mobile guidance, or a quadruped robot for dynamic display. The exact allocation depends on the current scene and robot status. This dynamic allocation is a practical expression of embodied intelligence in a multi-robot service environment.

4.3. Interaction and Feedback Loop

Interaction is the visible layer of the service process. Robots can interact with visitors through voice, visual recognition, gestures, guidance actions, and performance movements. The system fuses voice, vision, and action state information to understand visitor behavior and dynamically perceive interaction states.

Feedback closes the loop. The system continuously collects visitor interaction states and environmental feedback. If visitor needs or scene conditions change, the system adjusts interaction content and service strategies. This closed-loop design supports real-time collaborative service in complex cultural tourism scenarios.

5. Key Technology Research

5.1. Large Language Model-Driven Cultural Tourism Knowledge Service

The system uses a large language model as its core and combines it with a cultural tourism knowledge base to achieve visitor question understanding and intelligent question answering services. The cultural tourism knowledge base mainly includes scenic area information, historical culture, activity content, route recommendations, and public service information. These data are organized and stored in a structured manner.

When a visitor asks a question, the system first completes visitor intent recognition through a semantic parsing module. It then combines a retrieval-augmented generation mechanism to complete knowledge retrieval and context enhancement. Finally, the large language model generates the corresponding service content. This process enables intelligent question answering and interactive service in open scenarios.

In the context of embodied intelligence, the knowledge service is not only a conversational function. It is linked to robot roles and service actions. The generated content can become a guided explanation, a route recommendation, an activity introduction, or a performance cue. This linkage helps the system move from information provision to embodied service delivery.

5.2. Multimodal Human-Robot Interaction

To improve natural human-robot interaction in smart cultural tourism scenarios, the research integrates voice, vision, and action interaction into a multimodal collaborative interaction mechanism. In voice interaction, the system uses speech recognition and semantic understanding modules to achieve visitor voice question answering and real-time interaction. In visual perception, robots use visual recognition modules to achieve visitor state recognition and environmental perception. In action interaction, robots can provide interaction feedback through waving, guidance, and performance movements.

The system fuses voice, vision, and action state information to achieve visitor behavior understanding and dynamic interaction state perception. At the same time, the system dynamically adjusts robot interaction content and service strategies according to visitor interaction states. This supports real-time interactive service in complex cultural tourism scenarios.

Multimodal interaction is a key component of embodied intelligence because it connects perception to physical response. The robot does not merely process text. It observes, listens, moves, and reacts. This broader interaction capability is intended to make the service experience more natural and engaging for visitors.

5.3. Cross-Form Robot Collaborative Service Technology

To address the limited service capability of single robots and the insufficient adaptability of complex cultural tourism scenes, the research proposes a cross-form robot collaborative service mechanism to achieve multi-robot dynamic collaboration and scene linkage services. The system builds a unified multi-robot collaborative service framework based on differences in movement ability, environmental adaptability, and interaction capability among robots.

Humanoid robots are mainly responsible for tour guidance and visitor interaction. Quadruped robots are responsible for interactive parades and dynamic displays. Wheeled robots are responsible for mobile guidance and scene linkage. The system combines task requirements, robot operating states, and scene information to achieve dynamic task allocation and collaborative service among different robot forms. During service, the system can dynamically adjust robot collaboration strategies according to changes in visitor needs and environmental states, enabling real-time collaborative service in complex cultural tourism scenarios.

This cross-form collaboration is a distinctive feature of the proposed embodied intelligence system. Instead of relying on one robot to perform all functions, the architecture distributes tasks across multiple physical agents with complementary capabilities. The result is greater scene coverage and service flexibility.

5.4. Robot Group Control and Cultural Performance

To enhance interactivity and cultural experience in smart cultural tourism scenarios, the research integrates motion capture and robot motion choreography technologies to achieve multi-robot group performance and action interaction. The system collects human motion information through motion capture devices and uses a robot motion mapping mechanism to convert human motions into robot action sequences.

On this basis, the system combines group dance choreography and motion synchronization mechanisms to realize robot traditional Chinese group dance, Tai Chi demonstrations, and group action interaction. During performance, the system uses time synchronization and motion collaborative control mechanisms to achieve multi-robot motion coordination and group interaction. It also combines music rhythm and scene content to dynamically adjust robot action sequences, improving the观赏性 and interactivity of robot performances.

Within the logic of embodied intelligence, group control extends interaction from individual service to collective cultural expression. The robots become not only service providers but also participants in cultural presentation. This dual role helps connect smart tourism services with cultural experience.

6. Field Deployment at the Tianjin Five Great Avenues Begonia Festival

To verify the application effect of the proposed smart cultural tourism multi-robot interactive service system in real open scenarios, the research team deployed and validated the system at the Tianjin Five Great Avenues Begonia Festival. The deployment realized multiple service functions, including robotic tour guidance, visitor interaction, intelligent question answering, and robot group performance. The team also analyzed the system’s interaction capability, multi-robot collaboration capability, and scene adaptability.

6.1. Scenario Characteristics

The Tianjin Five Great Avenues Begonia Festival is a large urban cultural tourism event. It features a large number of visitors and complex interaction needs. These conditions place high requirements on the robot system’s interaction capability, real-time response capability, and adaptability to complex scenes.

In such an environment, embodied intelligence must operate under real-world constraints. The system must handle crowd movement, varied visitor questions, changing service demands, and the need for coordinated robot behavior. The field deployment was designed to test whether the proposed architecture could function under these conditions.

6.2. Robot Roles and Service Configuration

In response to the scenario characteristics, the research team constructed a multi-robot collaborative service system composed of humanoid robots, quadruped robots, and wheeled robots, combined with a cloud-based cultural tourism knowledge service platform for on-site deployment. Humanoid robots mainly undertook scenic area interpretation, visitor interaction, and cultural performance tasks. Quadruped robots were responsible for interactive parades and dynamic displays. Wheeled robots were responsible for mobile guidance and scene linkage services.

The system combined a large language model, multimodal interaction, and cross-form robot collaboration mechanisms to realize scenic area question answering, route recommendations, activity introductions, and interactive displays. It also combined robot group dance choreography and action interaction technologies to realize robot cultural interaction services such as traditional Chinese group dance and Tai Chi demonstrations, enhancing the interactivity and观赏性 of cultural tourism activities.

6.3. Cultural Performance and Visitor Engagement

The cultural performance function is one of the most visible applications of embodied intelligence at the festival. Robots performed coordinated movements based on motion capture and choreography. The performances were designed to reflect cultural themes and to interact with the festival atmosphere. By combining music rhythm and scene content, the system dynamically adjusted robot action sequences.

This performance capability extends the role of service robots beyond answering questions and guiding routes. It turns robots into active participants in cultural expression. For visitors, the combination of intelligent service and group performance may create a more memorable experience than a purely functional robot interaction.

7. System Functional Verification

To verify the service capability of the system in smart cultural tourism scenarios, the research team carried out application verification in areas including cultural tourism knowledge service, multimodal interaction, and multi-robot collaboration.

7.1. Cultural Tourism Knowledge Service Verification

In terms of cultural tourism knowledge service, the system was able to complete scenic area information queries, historical culture interpretation, activity content introductions, and route recommendations. It supported natural language question answering and multi-turn interaction in open scenarios. During actual application, the system could complete semantic understanding and knowledge retrieval in real time according to visitor questions, achieving intelligent cultural tourism knowledge service.

The combination of the large language model and cultural tourism knowledge base was central to this function. The retrieval-augmented generation mechanism helped ground responses in structured knowledge. This grounding is important for embodied intelligence because physical robots must provide reliable information in public service settings.

7.2. Multimodal Interaction Verification

In terms of multimodal interaction, robots were able to combine voice, vision, and action interaction to achieve visitor state recognition and real-time interaction. They dynamically adjusted interaction content and service methods according to visitor behavior. Compared with traditional single voice interaction modes, the proposed multimodal interaction mechanism improved the natural interaction capability of the robot system and the visitor participation experience.

This verification shows how embodied intelligence can enrich human-robot interaction. The robot’s ability to perceive visual and action cues, in addition to voice, allows it to respond more appropriately to the social context of the interaction.

7.3. Multi-Robot Collaboration Verification

In terms of multi-robot collaborative service, the system was able to dynamically schedule different forms of robots to complete tour guidance, interactive parades, and robot performance tasks according to scene needs. This achieved collaborative service capability in complex open scenarios. Compared with traditional single-robot service models, the multi-robot collaboration mechanism effectively improved the system’s scene coverage capability and service flexibility.

In addition, the system completed multiple rounds of robot interactive service and group performance tasks at the Begonia Festival. It realized various application functions such as visitor interaction, tour guidance, and cultural interaction, verifying the application feasibility of the system in real cultural tourism scenarios.

8. Application Results and Comparative Analysis

The actual application results show that the proposed smart cultural tourism multi-robot interactive service system can effectively improve intelligent service capability and visitor interaction experience in complex cultural tourism scenarios. The application effect comparison is presented in the table below.

Comparative application results at the Begonia Festival

Evaluation Indicator Traditional Service Model Application Results of the Proposed System Effect Description
Daily visitors served Approximately 100 visitors Approximately 300 visitors Multi-robot parallel service improves on-site service coverage capability.
Average response time Approximately 15 to 30 seconds Approximately 3 to 8 seconds Large language model knowledge service combined with preset interaction processes improves response efficiency.
Cultural tourism question answering success rate Approximately 75 percent More than 90 percent Cultural tourism knowledge base and large language model enhance semantic understanding capability.
Service coverage range Fixed points Multi-point linkage Humanoid, quadruped, and wheeled robots collaboratively cover guidance, parade, and performance scenarios.
Average visitor dwell time Approximately 2 to 3 minutes Approximately 5 to 10 minutes Group performance and multimodal interaction enhance visitor participation.

The comparative results indicate that the proposed system increased the daily number of visitors served from approximately 100 to approximately 300. The average response time decreased from approximately 15 to 30 seconds to approximately 3 to 8 seconds. The cultural tourism question answering success rate rose from approximately 75 percent to more than 90 percent. The service coverage range expanded from fixed points to multi-point linkage. The average visitor dwell time increased from approximately 2 to 3 minutes to approximately 5 to 10 minutes.

These findings suggest that embodied intelligence can contribute to practical improvements in service coverage, responsiveness, question answering accuracy, and visitor engagement. The results also show that multi-robot collaboration and multimodal interaction can work together to create a more dynamic service environment.

9. Significance for Embodied Intelligence Deployment

9.1. Operational Benefits for Smart Cultural Tourism

The deployment at the Begonia Festival demonstrates operational benefits for smart cultural tourism. The multi-robot architecture allowed different robot forms to serve different functions. Humanoid robots provided interpretation and interaction. Quadruped robots added dynamic display and parade capabilities. Wheeled robots supported mobile guidance and scene linkage. The cloud-based knowledge platform provided large language model inference and cultural tourism knowledge services.

For venue operators, this distributed service model may improve coverage and flexibility. For visitors, it may provide faster responses and more engaging interactions. The research team positions the system as a reference for applying embodied intelligence robots in smart scenic areas, public services, and urban cultural tourism.

9.2. Visitor Experience and Interaction Quality

Visitor experience is a central outcome of the proposed system. The combination of voice, vision, and action interaction supports more natural engagement. Visitors can ask questions, receive guidance, and participate in interactions with robots. The group performance function adds a cultural and entertainment dimension.

The increase in average visitor dwell time from approximately 2 to 3 minutes to approximately 5 to 10 minutes suggests that the system encouraged visitors to spend more time engaging with the service. This outcome is consistent with the design goal of using embodied intelligence to create richer and more interactive cultural tourism experiences.

9.3. Cultural Presentation and Public Engagement

The system also contributes to cultural presentation. Robot group dance and Tai Chi demonstrations combine motion capture, choreography, synchronization, and collaborative control. These performances are not merely technical demonstrations. They are designed to reflect cultural themes and interact with the festival atmosphere.

By turning robots into performers as well as service providers, the system expands the role of embodied intelligence in public cultural events. This dual function may help attract visitor attention, encourage participation, and strengthen the connection between technology and cultural content.

10. Limitations and Future Research Directions

10.1. Autonomous Collaboration in Complex Open Scenes

The research notes that future work will further study multi-robot autonomous collaboration and embodied intelligence interaction technologies for complex open scenarios. While the field deployment demonstrated feasibility, complex open environments continue to pose challenges for perception, decision-making, and coordination.

Future research may need to improve the robustness of environmental perception, the flexibility of task allocation, and the adaptability of service strategies. These improvements would support more autonomous collaboration among robot forms in crowded and dynamic public spaces.

10.2. Scalable Deployment Across Public Service Scenarios

The research also points toward exploring scalable applications of embodied intelligence robots in smart scenic areas, public services, and urban cultural tourism. Scaling from a festival deployment to broader public service scenarios requires attention to system reliability, knowledge base maintenance, robot fleet management, and user experience consistency.

As embodied intelligence technologies mature, multi-robot service systems may become more capable of supporting diverse tasks. These tasks could include visitor guidance, cultural interpretation, public information services, interactive performances, and coordinated event support. The proposed system provides an early reference for such developments.

10.3. Continued Integration of Large Language Models and Physical Robots

The integration of large language models with physical robots is a fast-moving area. In the proposed system, the large language model supports intent understanding, knowledge question answering, and service content generation. Multimodal perception connects the model to visitor states and environmental conditions. Cross-form robot collaboration translates service decisions into physical actions.

Further integration may improve the grounding of language model outputs in physical context. It may also enhance the ability of robots to coordinate across forms, adapt to changing scenes, and deliver culturally meaningful interactions. These directions align with the broader development of embodied intelligence.

11. Conclusion

The research addresses the limitations of traditional cultural tourism robot systems in complex open scenarios, including limited interaction capability, insufficient multi-robot collaboration, and single forms of interactive service. It proposes an embodied intelligence-based smart cultural tourism multi-robot interactive service system. The system integrates a large language model, multimodal interaction, and cross-form robot collaboration technologies. It constructs an integrated smart cultural tourism service architecture based on perception, understanding, interaction, and service.

The system realizes cultural tourism knowledge question answering, visitor interaction, and robot collaborative performance functions. It was deployed and validated at the Tianjin Five Great Avenues Begonia Festival. The results show that the proposed system can achieve multi-robot collaborative service and natural human-robot interaction in complex cultural tourism scenarios. It improves service capability, interactive experience, and cultural display effects in smart cultural tourism scenarios.

The research team states that future work will further study multi-robot autonomous collaboration and embodied intelligence interaction technologies for complex open scenarios. It will also explore scaled applications of embodied intelligence robots in smart scenic areas, public services, and urban cultural tourism. The proposed system offers a practical reference for the application of embodied intelligence in smart cultural tourism and related service domains.

12. References and Sources

  • Brohan, A., Brown, N., Carbajal, J., et al. RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control. 2023. DOI:10.48550/arXiv.2307.15818.
  • Driess, D., Xia, F., et al. PaLM-E: An embodied multimodal language model. arXiv:2303.03378, 2023.
  • Tao, Y., Wan, J., Wang, T., Xiong, Y., Wang, B., Zhang, W., Deng, C., Tao, Y., Yang, G., Wei, H. Building a new paradigm for embodied intelligence: A review of humanoid robot technology status and development trends. Journal of Mechanical Engineering, 2025, 61(15), 121-147.
Scroll to Top