A newly published systematic review led by researchers at Beihang University, together with collaborators from the Chinese Institute of Electronics, the Hangzhou International Innovation Institute of Beihang University and Zhejiang University, offers one of the first dedicated technology and industry roadmaps for wheeled humanoid robots driven by embodied intelligence. The work argues that the convergence of wheeled mobility, dual-arm dexterous manipulation and large-scale learning systems is turning these platforms into one of the most practical physical carriers for embodied intelligence, and that the field is now entering a decisive window between technical validation and large-scale engineering deployment.
Unlike purely virtual artificial intelligence, which processes information inside digital spaces, embodied intelligence requires a physical body that can interact with the real world, adapt to unstructured environments and execute long chains of complex tasks. Wheeled humanoid robots combine the high mobility of a wheeled chassis with the collaborative operation capability of two independent arms, and they unify mobility and manipulation inside a single perception, planning and control framework. According to the review, this combination gives them a clear engineering edge in factories, warehouses, equipment maintenance, component assembly and service scenarios.

The review draws a deliberate boundary around the category. A mobile platform carrying a single manipulator is limited in flexibility and collaboration. A fixed dual-arm robot lacks workspace reach and environmental adaptability. A bipedal humanoid robot still faces open questions in stability, energy consumption, cost and engineering maturity. Wheeled humanoid robots, by contrast, are positioned as the configuration that best balances mobility, manipulation capability and deployability, providing a practical foundation for industrial upgrading and the physical implementation of embodied intelligence.
1. A Category Defined by Mobility and Manipulation Integration
The researchers describe wheeled humanoid robots as systems built on a wheeled mobile base, equipped with two robotic arms that can work independently or cooperatively, and governed by a unified perception, planning and control architecture that enables integrated mobility-and-manipulation tasks. This definition distinguishes them from three neighboring families of systems and establishes a clear technical scope for both academic research and industrial product planning.
The motivation for this architecture is closely tied to the evolution of embodied intelligence itself. As intelligent robots move from programmed execution toward autonomous interactive operation in dynamic and open environments, the ability to handle long-horizon tasks becomes the central performance metric. Embodied intelligence in this context is not only about perception accuracy or control bandwidth; it is about the capacity of a physical agent to understand a task, decompose it, move to the required location, manipulate the required object and recover from failure. Wheeled humanoid robots are increasingly viewed as the platform where all of these embodied intelligence capabilities must be integrated simultaneously.
The review notes that existing literature remains fragmented. Most studies optimize a single technical module, and most surveys focus on general humanoid robots or generic mobile manipulators. Dedicated summaries of wheeled humanoid robots, including engineering bottlenecks and industrial development paths, have been comparatively scarce. The new work is intended to fill that gap by treating the platform as a complete system rather than a collection of subsystems.
2. International and Domestic Development Landscape
The review compares the development trajectories of wheeled humanoid robots across different regions and finds two distinct but complementary emphases: internationally, work has concentrated on foundational methods and system architecture innovation; domestically, the emphasis has shifted toward system integration, stability assurance and scenario-oriented complex manipulation capability, with rapid product iteration and early commercial pilots.
2.1 International research: system-level verification and architecture innovation
Internationally, the field has been driven mainly by university research groups, with a smaller number of start-ups conducting system-level verification and prototype exploration. Early work focused on simply stacking a mobile base and one or more manipulators. As the need for coordinated control became clearer, research evolved toward tightly coupled architectures in which mobility, manipulation and perception are treated as one integrated system.
An early landmark is the two-arm mobile manipulator developed at the Technical University of Munich, which systematically studied unified modeling, whole-body motion planning and stability control for a wheeled chassis working together with two arms. It is widely regarded as one of the first international benchmarks for treating mobile manipulation as a single problem rather than two separate ones. Another representative platform is the Rollin’ Justin mobile dual-arm system from the German Aerospace Center, which combines a wheeled mobile base with humanoid dual-arm upper limbs and targets autonomous mobile manipulation in complex environments, with research focusing on whole-body coordination, hybrid force-position control and safe mobile manipulation under constraints.
More recently, the University of Hamburg has demonstrated notable strengths in multimodal perception fusion and embodied cognitive learning. Its work on dual-arm coordination and whole-body motion generation includes a heuristic inverse kinematics framework based on multi-objective cost functions, which effectively addresses dual-arm coordination and collision avoidance under high redundancy, and has been validated on wheeled humanoid platforms performing complex dual-arm handover and fine sorting. Rather than relying purely on model-driven control, this line of research emphasizes complementary multi-sensor information and sim-to-real transfer learning strategies, enabling efficient coupling between the wheeled chassis and the arms and adaptive decision-making in dynamic and uncertain environments.
On the industrial side, start-ups such as the United Kingdom-based developer of the HMND 01 Alpha dual-arm mobile robot represent recent attempts to build general-purpose mobile manipulation platforms in the embodied intelligence era. These platforms emphasize dual-arm collaboration and multimodal perception fusion and target multi-task execution and general manipulation validation, but they largely remain at the prototype demonstration and technical verification stage. Overall, international research leans toward foundational methods and system architecture, while mature commercial products and large-scale applications remain relatively limited, reflecting the high barrier of technical integration and reliability verification in wheeled dual-arm systems.
2.2 Domestic development: engineering integration and rapid product iteration
In contrast, domestic development has advanced rapidly, showing a pattern in which technical research and engineered products proceed in parallel. The emphasis is on system integration, stability assurance and the construction of complex manipulation capabilities oriented toward real scenarios, laying the groundwork for later application deployment.
Several representative products illustrate this direction. UBTech has introduced a full-size general-purpose wheeled humanoid robot equipped with a fourth-generation dexterous hand capable of sub-millimeter precision operation and integrated learning-based motion control, enabling adaptive human-robot safety in complex environments. Paxini’s TORA DOUBLE ONE has been specifically designed for dual-arm collaboration, force perception and complex interactive manipulation, using a dual steering-wheel AGV chassis with a foldable structure and a dual-modal dexterous hand with nearly 2,000 tactile sensing units, suited to logistics and warehousing scenarios. Galaxy General’s GALBOT G1 focuses on general mobile manipulation tasks, emphasizing dual-arm coordination and autonomous perception, and is equipped with a cross-embodiment navigation foundation model that enables autonomous operations such as material sorting in industrial manufacturing scenarios.
Tencent Robotics X has introduced a robot with a four-leg wheel-legged hybrid design, large-area tactile skin, multi-finger dexterous hands and safe physical human-robot interaction. Other developers have concentrated on embodied dual-arm robots built around self-developed operating systems, emphasizing the unification of mobility, manipulation and perception. A start-up spun out of Tsinghua University has released a wheeled humanoid robot with 23 degrees of freedom, a 360-degree LiDAR and multimodal sensors, positioning the platform as a multifunctional base for embodied intelligence developers at home and abroad. Industrial robotics companies have strengthened autonomous navigation and dual-arm cooperative operation for industrial scenarios, forming a complete technology chain that spans structural design, system integration and adaptation to complex environments.
Universities have laid a solid technical foundation. Leveraging intelligent robot control and embedded mechatronics, research teams have achieved real-time multi-sensor fusion and high-precision motion control in system integration, while focusing on structural topology optimization and dynamic balancing algorithms to improve adaptability in unstructured environments. Beihang University’s embodied intelligence robotics institute has independently developed a wheeled humanoid robot platform capable of heavy-load handling tasks, supporting embodied intelligence teaching and research. The humanoid robot innovation center at Zhejiang University has developed a wheeled humanoid robot specialized for ultra-precision operation at the 0.03 mm level, which has already been applied at scale in the production lines of leading enterprises in the electronics and automotive sectors.
The following table summarizes representative domestic wheeled humanoid robot products and their physical, perceptual and functional characteristics, as compiled in the review.
| Developer | Release | Product | Physical and motion specifications | Perception and operation capability | Function and target scenarios |
|---|---|---|---|---|---|
| Galaxy General | 2024.6 | Galbot G1 | 173 cm / 85 kg / 43 DOF; omnidirectional wheeled leg design with 360-degree movement | Adapts to different terrain; generalized grasping and environment adaptation | Industrial and service manipulation tasks |
| Tencent | 2024.9 | Xiao Wu | 140-180 cm / 80 kg / 29 DOF; multimodal wheel-legged configuration; dual-arm carrying load 50 kg | Supports autonomous folding | Complex terrain mobility and heavy-load transport |
| Paxini | 2025.8 | TORA DOUBLE ONE | 167-169 cm / 70-73 kg / 20-22 DOF; dual-wheel AGV and four-wheel-drive folding chassis | Nearly 2,000 tactile sensing units; hand payload 6.5 kg | Fine tactile operation, logistics and warehousing |
| UBTech | 2025.8 | Cruzr S2 | 176 cm / 185 kg / 44 DOF; wheeled chassis at 2 m/s; transport payload 15 kg | Full-space handling from 0 to 1.8 m | Guided tours, medical transfer, production assembly |
| Unitree | 2025.11 | G1-D Flagship | 126-168 cm / 80 kg / 19 DOF; wheeled chassis at up to 1.5 m/s; single-arm payload 3 kg; 6 h endurance | Jetson Orin NX at 100 TOPS | Collaborative dual-arm tasks, complex environment mobility |
| RealMan | 2025.8 | RMC-AIDAL | 170 cm / 90 kg / 12 DOF | Binocular D435 vision plus RGB monitoring; two RM65-B-V ultra-light arms | Collaborative dual-arm tasks |
| AgiBot | 2025.10 | Genie G2 | 185 kg / 50 DOF with 26 active joints; high-performance motion joints | High-precision torque sensors; multimodal language interaction | Industry, logistics, guidance |
| Robot Era | 2025.6 | Q5 | 165 cm / 65 kg / 44 DOF; compact wheeled chassis | LiDAR and vision fusion navigation; single-arm payload 10 kg | Autonomous obstacle avoidance in narrow passages |
| Dobot | 2025.3 | ATOM-W | 153 cm / 62 kg / 41 DOF; wheeled chassis at 1.5 m/s; repeat positioning accuracy of ±0.05 mm | 7-DOF industrial-grade collaborative arms | Fine operations such as wafer transfer |
| Keenon Robotics | 2025.3 | XMAN-R1 | Humanoid structure with 360-degree high-precision perception | 11 multimodal sensors plus a large language model | Restaurant and service interaction loop covering ordering, delivery and table clearing |
| DexForce | 2025.7 | W1 Pro | 34 power units; highly humanoid structure | Pure-vision spatial intelligence sensors | Precise control and real-time perception; coffee making and interactive tasks |
| SEER Robotics | 2025.10 | X1-PRO | 40 DOF; omnidirectional chassis with 45-degree bending structure; single-arm payload 6 kg and instantaneous dual-arm payload 16 kg | Integrated industrial-grade perception | Logistics sorting, loading and unloading, collaborative handling |
| Youibot | 2025.8 | Xunxiao | Industrial-grade mobile chassis | One-brain, multiple-form industrial embodied intelligence model | Semiconductor precision manufacturing, power inspection, disaster prevention |
3. The Five Core Technology Systems Behind Embodied Intelligence
The review organizes the technology stack of wheeled humanoid robots into five interconnected systems: body hardware and supporting architecture; mobile chassis environmental perception and motion planning; task understanding and dual-arm dexterous manipulation; data acquisition and simulation training for complex scenarios; and embodied intelligence control paradigms with sim-to-real transfer. Together these systems support mobile manipulation across diverse scenarios and push robots from programmed execution tools toward agents with autonomous perception, decision-making and learning capability.
3.1 Body hardware, components and core architecture
The physical design of a wheeled humanoid robot is the foundation of both motion and manipulation. Its core hardware architecture is typically modular and heterogeneous, comprising the mobile chassis, a lifting mechanism, dual arms, dexterous hands and end effectors, together with supporting systems such as motion control, communication, operating systems and middleware. The review observes that modularity and lightweight design have become mainstream directions, yet a unified design theory and method for whole-machine dynamics optimization under mobility-manipulation coupling, and for reliable multi-sensor system integration, has not yet been established.
Within the physical configuration, the mobile chassis and the lifting mechanism together determine the robot’s operating radius and vertical working space. Chassis selection must balance flexibility, payload and environmental adaptability. Current mainstream solutions show a trend from simple constrained motion toward omnidirectional high-dynamics capability. Common chassis types include Mecanum wheels, steering-wheel configurations, differential drives and omnidirectional wheels, each offering a differentiated trade-off among omnidirectional mobility, load capacity, control complexity and environmental adaptability. To address the large vertical span required by dual-arm operation, lifting mechanisms have become a key component connecting the chassis and the waist, enabling the robot to raise and lower its upper body. Linear guide rail mechanisms offer good rigidity and simple control but are limited by slider stroke and compactness, while scissor-folding mechanisms are structurally compact and can achieve a high lifting ratio, at the cost of stricter manufacturing precision and lateral force resistance requirements. This combined mobile-and-lifting architecture allows wheeled humanoid robots to combine efficient movement on indoor flat floors with flexible reach across complex height environments.
The arm configuration of a wheeled humanoid robot is centered on balancing motion flexibility, structural stiffness and spatial adaptability. Common configurations include serial, parallel and hybrid serial-parallel types. Hybrid configurations, which combine the large workspace of serial mechanisms with the high stiffness of parallel mechanisms, have become an important direction for balancing operating range and load capacity. Modular design is a key approach to configuration optimization, enabling flexible structural expansion and easier maintenance through standardized joint and link units, while lightweight materials reduce energy consumption and improve motion response speed.
Underlying arm control targets precision and compliance. Main methods include position control, force and torque control and impedance control, all of which use closed-loop feedback to correct joint motion errors and ensure accurate trajectory tracking. Force control and impedance control enable arms to adapt when interacting with environments or humans, dynamically adjusting motion parameters based on sensed contact force to achieve compliant operation.
As the direct physical interface between the arm and the work object, the end effector is designed around dexterity, generality and sensing integration. Humanoid multi-finger dexterous hands, with anthropomorphic structural design and multi-dimensional manipulation capability, have become the core carrier for complex operation. They use multi-finger, multi-joint bionic structures and combine flexible joints with tendon and gear hybrid transmission mechanisms, allowing them to adapt to objects of different shapes, materials and stiffness while balancing grasping stability and operational safety. Deep integration of perception and control is the defining feature that distinguishes dexterous hands from conventional end effectors. High-density tactile, force and position sensors integrated in fingertips, finger pads and the palm collect contact pressure, contact pose and surface topography in real time, providing high-precision, high-real-time feedback for low-level control. Control logic is deeply coordinated with low-level arm control, enabling fine adjustment of grasping force and smooth motion transitions, which is essential for precise assembly and in-hand manipulation. Current development directions focus on improving perception accuracy, strengthening multi-task adaptability and optimizing whole-body cooperative control with the arms and mobile chassis.
Domestic dexterous hand technology has advanced rapidly. Multi-finger bionic hands using tendon-driven structures and multimodal tactile sensing systems can achieve high-DOF grasping and in-hand manipulation, showing strong potential in industrial assembly and service robotics. Modular dexterous hands with high-precision servo drives provide multi-finger cooperative grasping and are widely used in research and education. Continued optimization of structural design, tactile perception and control algorithms in these products provides important technical support for complex manipulation tasks on wheeled humanoid robots.
As the critical infrastructure connecting the mechanical body to upper-level intelligent algorithms, the hardware structure and software architecture together form the core computing and control system. The system must provide high computing power for multimodal perception processing and intelligent decision-making, while simultaneously meeting the strict real-time requirements of chassis motion control and arm servo drives. As a result, modern wheeled humanoid robots generally adopt a layered architecture in which hardware computing and software platforms are co-designed to allow efficient coordinated operation of perception, planning and control tasks. To handle complex control chains and multi-source data processing, research and engineering practice typically adopt a layered heterogeneous computing architecture, using a unified software platform to schedule heterogeneous hardware resources and abstract functionality.
In terms of hardware computing architecture, systems usually adopt a brain-cerebellum layered structure. The brain portion carries high-performance CPU or GPU computing platforms to execute environment perception, semantic understanding, task planning and vision-language models, while the cerebellum portion consists of real-time embedded or industrial controllers responsible for chassis motion control, joint servo driving and high-frequency sensor data acquisition. This heterogeneous structure preserves millisecond-level control response while providing sufficient computing resources for complex intelligent algorithms, enabling coordinated operation of mobility, manipulation and perception.
At the software architecture level, robot operating systems and middleware form the essential platform connecting hardware devices and upper-level applications. Current wheeled humanoid robots generally adopt a dual-system architecture in which a general-purpose operating system and a real-time operating system work together: the general-purpose system handles upper-level algorithm execution, data processing and software ecosystem support, while the real-time system focuses on low-level motion control and real-time task scheduling. Robot middleware such as ROS uses node-based communication to enable data exchange and functional coordination among modules, and its publish-subscribe mechanism creates high-bandwidth, low-latency data channels that greatly reduce the development complexity of multi-sensor fusion, dual-arm cooperative control and task scheduling systems. The rich function packages and development tools in the ROS ecosystem allow SLAM, motion planning and vision processing modules to be integrated quickly.
At the device communication level, industrial fieldbuses and real-time communication protocols provide critical support for stable hardware operation. Buses such as EtherCAT and CAN enable unified addressing and high-speed communication with chassis drivers, joint servo systems and various sensors, and distributed clock mechanisms achieve system-level time synchronization that guarantees coordinated operation of multiple actuators in complex tasks. To reduce system integration complexity and improve software scalability, modern robot systems also build hardware abstraction layers that uniformly encapsulate low-level hardware resources. Standardized interfaces shield differences among hardware platforms, communication protocols and mechanical structures, packaging chassis mobility, arm manipulation and sensing capabilities into standardized software interfaces or functional modules so that upper-level algorithms can be developed and deployed in a unified software environment. This software-centric design philosophy gives the system strong modular expansion capability, facilitates rapid integration of heterogeneous hardware platforms and provides a stable foundation for deploying and iterating complex intelligent algorithms.
3.2 Mobile chassis environmental perception and motion planning
Environmental perception and motion planning for the mobile chassis are the foundation of autonomous mobile operation, addressing environment modeling, pose estimation, path generation and trajectory tracking. Compared with conventional indoor mobile robots, wheeled humanoid robots face more dynamic and open application environments, imposing higher requirements on perception robustness, planning real-time performance and control precision. Existing navigation methods can be summarized from two perspectives: path planning strategy and control implementation. They include graph search planning, sampling-based planning, local planning, optimization-based control, and the learning-based and semantic navigation methods that have emerged in recent years.
| Method category | Representative algorithms | Advantages | Limitations | Typical application |
|---|---|---|---|---|
| Graph search planning | A*, D* Lite | Good stability and strong interpretability | Weaker adaptability to dynamic environments | Global path planning |
| Sampling-based planning | RRT, RRT*, PRM | Suitable for complex space search | Path smoothness is generally limited | Feasible path generation |
| Local planning | DWA, TEB | Strong real-time performance | Prone to local optima | Dynamic obstacle avoidance |
| Optimization-based control | MPC | Can handle dynamic constraints | High computational cost | Trajectory tracking control |
| Learning-based navigation | Deep reinforcement learning | Strong adaptability in complex environments | High training cost and insufficient stability | Navigation in unknown environments |
| Semantic navigation | Semantic navigation frameworks | Supports goal-level navigation | Depends on semantic perception accuracy | Service scenario navigation |
In environmental perception, localization and mapping, a single sensor can hardly balance accuracy, robustness and adaptability simultaneously. Mainstream solutions therefore adopt multi-sensor fusion architectures combining LiDAR, vision and inertial measurement units. LiDAR is suited to acquiring stable geometric structure information, vision sensors provide texture and semantic information, and IMUs enhance short-term motion estimation stability. Through tightly coupled fusion of multiple sources, the mobile chassis can achieve more stable localization and mapping, and map representation is gradually evolving from purely geometric maps toward maps that incorporate scene semantics.
In path planning and motion control, existing methods typically use a hierarchical structure combining global and local planning. Global planning generates reachable paths based on a prior map, commonly using graph search algorithms such as A* and D* Lite, or sampling-based algorithms such as RRT, RRT* and PRM. Local planning addresses dynamic obstacles and local passage constraints, often using the dynamic window approach or the time elastic band method for real-time obstacle avoidance. Graph search methods offer good stability, sampling-based planning is better suited to complex spaces, and local planning provides strong real-time performance in dynamic scenarios. At the trajectory execution level, the mobile chassis needs closed-loop control to track planned paths stably. Model predictive control can explicitly consider kinematic and dynamic constraints within a prediction horizon and is therefore widely used for trajectory tracking control of mobile chassis. Building on this, research has introduced semantic navigation methods that identify doors, tabletops, corridors and other semantic targets to assist path decision-making, improving navigation efficiency in service and indoor operation scenarios. In addition, navigation methods based on deep reinforcement learning are increasingly used for autonomous obstacle avoidance and policy learning in unknown environments, showing adaptability in complex scenarios, although engineering stability and training cost still require further study.
3.3 Task understanding and dual-arm dexterous manipulation
Compared with single-arm operation, the core problem for wheeled humanoid robots is no longer a single grasp or trajectory generation, but task-centric dual-arm collaborative understanding and execution. Dual-arm tasks usually require explicit operating objectives, constraints and execution order, including the geometric and functional attributes of target objects, safety and collision-avoidance constraints during operation, and the temporal relationships of multi-step tasks. On this basis, research generally classifies dual-arm collaboration into several typical modes: master-slave collaboration, in which one arm performs the main operation while the other provides support or constraint; symmetric collaboration, which emphasizes equivalence in load distribution and motion execution and is often used for cooperative carrying and symmetric assembly; and hybrid collaboration, which dynamically switches arm roles according to task stage and is closer to real complex task requirements. Such task representations provide high-level semantic constraints for subsequent motion planning and control.
In dual-arm cooperative motion planning, existing research builds on three main technical routes. The first is inverse-kinematics-based planning, which jointly solves dual-arm joint configurations to satisfy end-effector poses and collaboration constraints, offering high computational efficiency and easy integration. The second is sampling-based planning, which introduces dual-arm coupling constraints in a high-dimensional configuration space to improve the search for feasible solutions in complex environments. The third is optimization-based or task-space constraint representation, which explicitly models grasp pose families, tool coordinate frames and constrained motion trajectories as optimization objectives or constraints, making it more suitable for fine assembly and tool operation tasks. These methods form the technical basis of current dual-arm cooperative planning.
On the perception side and at the system integration level, dual-arm dexterous manipulation imposes strict requirements on the dimensionality and precision of multimodal information fusion. A representative example in multimodal fusion and advanced dexterous operation is a robot body interconnection architecture jointly introduced by Tencent Robotics X and the Futian Laboratory, which builds a full-stack communication system from the physical layer to the application layer, supports high-precision clock synchronization and real-time data transmission, and integrates seamlessly into mainstream robot development ecosystems, enabling advanced functions such as multi-joint coordination and efficient over-the-air upgrades. In addition, an open-source dual-arm platform led by technology teams provides an open hardware demonstration for contact-intensive physical AI research; its body design balances high dynamic response with compliant safety, accelerating the validation loop for vision-language-action models, offline reinforcement learning and sim-to-real strategies in complex dual-arm collaborative tasks.
3.4 Data acquisition and simulation training for complex scenarios
High-quality data is the cornerstone of embodied intelligence evolution, and teleoperation is a key pathway for acquiring high-fidelity physical interaction data. For long-horizon and contact-intensive tasks, a low-cost master-slave isomorphic mapping system has been shown to effectively reduce the complexity of inverse kinematics solving and improve operation precision compared with virtual reality schemes that lack force feedback. To break through laboratory scene limitations, a portable in-the-wild data collection system combining a handheld fisheye camera with SLAM technology has been proposed. With the development of mobile manipulation, teleoperation is evolving toward whole-body coordination of wheeled humanoid robots. A mobile dual-arm system with whole-body teleoperation enables coordinated input of the chassis and both arms, solving the safety control problem under nonholonomic constraints while maintaining low latency and a strong sense of presence, and providing an efficient technical means for interaction data collection in complex dynamic environments.
| Technology category | Core characteristics | Advantages | Limitations | Core applicable scenarios |
|---|---|---|---|---|
| VR teleoperation | Head-mounted visual immersion with 6-DOF controller pose tracking | Intuitive spatial perception, safe uncoupled operation, accessible consumer hardware | Lacks force feedback, difficult human-robot motion retargeting, prone to dizziness during long sessions | Long-range navigation, non-contact pick-and-place tasks |
| Motion capture | Optical markers or purely visual contactless human motion tracking | Natural and unconstrained operation, supports whole-body motion capture, adaptable to complex dexterous gestures | Sensitive to occlusion, accuracy fluctuates in purely visual schemes, inherent human-robot mapping errors | Dexterous hand fine manipulation, complex gesture interaction tasks |
| Exoskeleton and master-slave arms | Mechanical coupling with isomorphic master-slave mapping | High acquisition precision, native force feedback, data usable without complex retargeting | High hardware cost, limited operating space, physically demanding over long sessions | Precision assembly, contact-intensive manipulation tasks |
| Force feedback devices | Desktop pen-type haptic interaction devices | High force rendering precision, excellent micro-manipulation control | End-point control only, limited workspace, steep learning curve | Remote surgery, micro-manipulation, rigid surface probing |
| Kinesthetic teaching | Direct dragging of the robot end effector or joints | No additional hardware cost, error-free pose mapping | Force sensor data susceptible to interference, prone to visual occlusion, only suitable for lightweight collaborative arms | Simple trajectory reproduction, trajectory planning tasks without visual input |
To maximize data value, it is necessary to establish strict time-synchronized collection specifications covering vision, force, joint states and odometry. The core value of data lies not only in successful demonstrations but also in recording failures and recoveries. Closed-loop data that includes anomaly recovery has been confirmed to significantly improve model robustness. In addition, segmenting, cleaning and detecting contact events in long-chain trajectories can effectively improve the efficiency of offline reinforcement learning. Large-scale cross-embodiment projects further show that generalizable data with task diversity, environment diversity and coverage of various disturbances, combined with unified action space annotation, is the key to cross-scenario and cross-platform skill transfer.
To address the scarcity of physical data, high-fidelity simulation based on physics engines has become a mainstream solution. GPU-parallel computing has greatly improved sampling efficiency and supports large-scale domain randomization to cover the physical parameter distribution of the real world. One widely used manipulation benchmark enhances skill generalization through fine-grained physical property modeling and multi-sensor simulation. Recently, scene construction techniques combined with generative AI have become a new trend: large language models can automatically generate semantically rich 3D scenes, solving the time-consuming problem of traditional simulation scene construction and providing large-scale training environments with physical realism and visual diversity for mobile manipulation, strongly promoting zero-shot sim-to-real transfer.
| Simulation platform | Core engine | Core characteristics | Advantages | Limitations | Suitable task types |
|---|---|---|---|---|---|
| SAPIEN | Customized PhysX | Part-level physical interaction of articulated objects | High physical interaction fidelity, mature open-source ecosystem, rich datasets | Limited visual rendering capability, insufficient GPU parallelism in early versions | Precision manipulation, tool use, part assembly |
| Isaac Sim | PhysX 5 on Omniverse | GPU-parallel large-scale simulation | Leading training efficiency, realistic visual rendering, well-integrated ROS ecosystem | High hardware threshold, closed core engine, tiny contact simulation prone to penetration artifacts | Motion control, large-scale reinforcement learning, general grasping |
| Ai2-THOR | Unity | Semantic interactive household environments | Rich scene semantics, high visual fidelity, active community ecosystem | Simplified physics models, relatively low frame rates | Long-horizon task planning, visual navigation |
| Habitat | Bullet | High-frame-rate navigation simulation based on real scene scans | Extremely high single-machine frame rates, scenes built from real house scans | Weak object interaction capability, mainly rigid body simulation, no soft body or fluid simulation | Visual navigation, basic mobile manipulation |
| RFUniverse | Unity | Multi-physics full-modal simulation | Supports fluid, soft body and gas multi-physics simulation, well-developed VR interfaces | High computational cost for multi-physics, steep API learning curve | Complex physical interaction, flexible object manipulation, daily household tasks |
| ThreeDWorld | Unity | Multimodal perception physics simulation | Unique physics-based audio rendering, flexible procedural scene generation | Low adoption in robot manipulation, non-mainstream interfaces | Audio-visual fusion learning, physical commonsense reasoning |
Current simulation platforms show a parallel trend of specialization and generalization. Platforms oriented toward precision manipulation continue to deepen physical interaction fidelity, while platforms oriented toward large-scale reinforcement learning focus on computing efficiency and scene generalization, with the technical boundaries of the two gradually converging.
3.5 Embodied intelligence control paradigms and sim-to-real transfer
In mobile manipulation tasks oriented toward complex dynamic environments, the core problem of the control system is not only low-level motion control precision, but how to achieve a stable closed loop from semantic objectives to executable actions under long-horizon tasks, perception uncertainty and contact interaction constraints. The review organizes this area around three threads: control paradigms, intelligent decision-making tools, and training and deployment mechanisms.
Hierarchical control paradigms typically decompose robot tasks into high-level logical planning, mid-level skill primitives and low-level motion control, achieving dimensionality reduction and modular solution of control problems. The high-level planning layer generates physically feasible sub-goal sequences based on task objectives and environmental states, and can constrain and screen candidate actions using evaluation functions such as manipulability, safety distance and energy consumption. The mid-level module builds a skill library covering parametric primitives such as grasping, pushing, handing over and opening and closing, providing semantic continuity between task planning and motion control. Low-level control handles high-frequency closed-loop execution of the chassis, arms and end effectors, ensuring trajectory precision and contact compliance through position control, force and torque control or impedance control. The system can also integrate perception-feedback-based anomaly monitoring and recovery mechanisms, triggering local replanning when target displacement, grasp failure or collision risk is detected, forming a working logic with closed-loop self-correction capability.
In contrast to hierarchical control, which relies on explicit task decomposition and intermediate skill representation, end-to-end control is a data-driven paradigm whose core is to use deep neural networks to establish a direct mapping from raw perceptual input to action commands. In imitation learning, diffusion policy models action generation as a conditional diffusion process, better representing multimodal action distributions and alleviating the covariate shift problem of traditional regression models in complex manipulation tasks. In reinforcement learning, policies can explore optimal solutions by maximizing cumulative reward, and can combine large-scale offline data pretraining with online fine-tuning to improve the generalization of visual grasping and mobile manipulation policies. The advantage of end-to-end methods is that they can implicitly learn the associations among object properties, contact modes and action sequences from data, but their sample efficiency, interpretability and safety verifiability remain limited, so engineering deployment usually requires combination with high-fidelity simulation, constraint checking and safety control modules.
Beyond these control paradigms, multimodal large models are best positioned as tools for task understanding and high-level decision-making. Some frameworks place large models primarily in the role of high-level reasoners that translate ambiguous natural language instructions into structured planning prompts rather than directly outputting low-level continuous control signals. Other frameworks use vision-language-action models to discretize robot actions into tokens that can be predicted autoregressively by a large model, improving generalization in open-vocabulary tasks. Still other research uses large language models to generate code or build 3D value maps, combining semantic understanding with conventional motion planning. The main role of large models in wheeled humanoid robot systems is therefore to provide task decomposition, goal constraints and skill scheduling for hierarchical control, or to provide semantic conditions and prior knowledge for end-to-end policies. Given that large models may hallucinate, lack physical commonsense and have unclear safety boundaries, practical deployment still requires feasibility verification, collision detection and safety constraint mechanisms.
Sim-to-real transfer is the core technical means of bridging the gap between simulation and real physical scenarios and reducing the cost of acquiring embodied intelligence data. Domain randomization applies large-range perturbations to physical properties and visual parameters in simulation, forcing policy networks to extract robust features, and has been successfully applied to complex dexterous hand manipulation tasks. However, offline domain randomization alone cannot cover factors that change over time in real systems, such as wheel-ground contact, load disturbances, joint friction and sensor noise. Therefore, between domain randomization and online system identification, a transition stage of parameter adaptation and latent variable estimation is needed: simulation randomization expands the dynamics distribution encountered during policy training, while the robot’s proprioceptive history, action inputs and state responses are used to estimate implicit parameters of the current environment or body dynamics, which serve as conditioning inputs for the policy, achieving a connection from offline robust training to online adaptive control. Within this framework, rapid motor adaptation methods can be understood as a representative online adaptive path: an adaptation module estimates environmental latent variables such as ground friction, load variation and dynamic deviation from short-term proprioceptive history, then drives the base policy to adjust action output in real time, improving system adaptability to uncertain physical changes.
In summary, the embodied intelligence control technology of wheeled humanoid robots can be understood at three levels: control paradigms, intelligent decision-making tools, and training and deployment mechanisms. Hierarchical control and end-to-end control represent two main paradigms, explicit task decomposition and data-driven mapping. Multimodal large models primarily handle high-level decision functions such as task understanding, semantic reasoning and skill scheduling, and provide goal constraints and prior knowledge for control policies. Sim-to-real transfer supports the training, validation and engineering deployment of control policies through domain randomization, parameter adaptation, online system identification and real-world data feedback. These technologies occupy different levels in the perception-decision-planning-control-execution chain, and their coordinated integration is a key direction for improving long-horizon task execution, open-scenario generalization and engineering deployment reliability.
4. Typical Application Scenarios
The review identifies four main application domains where wheeled humanoid robots are being piloted or deployed: industrial manufacturing, commercial services, household applications and special-environment operations. In each case, the value proposition of embodied intelligence lies in combining autonomous mobility with dexterous manipulation in environments that were previously designed for humans.
4.1 Industrial manufacturing
In industrial manufacturing, wheeled humanoid robots are gradually being applied to material handling, precision assembly and equipment inspection in flexible factories. Compared with fixed-position robots, they can operate across multiple workstations and production rhythms thanks to their wheeled chassis, while the dual-arm structure supports cooperative grasping, alignment and assembly and tool operation. This form factor is well suited to manufacturing modes characterized by fast production rhythms, frequent process changes and small-batch, multi-variety output, helping to address the limitations of traditional automated production lines in flexibility and scalability.
Specifically, new energy vehicles and consumer electronics manufacturing are among the more typical early deployment scenarios. In power battery module assembly, wire harness organization and precision component assembly, dual-arm collaboration can significantly improve operational stability and assembly accuracy. In equipment inspection and maintenance, mobile manipulation capability allows the robot to complete autonomous operations from detection to intervention.
Several enterprises have already achieved early industrial deployment of wheeled humanoid robots. One manufacturer’s wheeled humanoid robot has been adapted to a full range of industrial scenarios and achieved large-scale commercial deployment, with multiple units flexibly deployed in core production workshops to significantly improve production line automation and flexible production efficiency. Another company’s wheeled humanoid robot has achieved large-scale industrial deployment, with nearly one hundred units deployed in core workshops to enable continuous operation and substantially reduce manual intervention costs. A third platform offers long endurance, large payload and efficient stability, combining data-driven industrial embodied intelligence models with an industrial software platform to closely match industrial production rhythms. A heavy-payload product designed for industrial scenarios addresses high-intensity manual operations such as cross-station transfer of heavy materials and collaborative assembly, achieving fully autonomous operation and providing a new technical path for flexible upgrading of heavy manufacturing. Additional systems have been deployed in precision assembly scenarios across multiple fields, efficiently completing various collaborative tasks.
In special-environment service, such as substation and large facility inspection, dual-arm mobile manipulation systems can perform inspections while executing simple operations or emergency handling. From an industrial practice perspective, wheeled humanoid robots are restructuring the automation model of industrial manufacturing. They break the limitations of traditional fixed-position robots and, with one machine serving multiple functions and covering the whole site, adapt to the flexible production needs of small-batch, multi-variety, fast-iterating manufacturing, providing a replicable and scalable engineering path for embodied intelligence technology in industrial manufacturing.
4.2 Commercial services
In commercial services, the capabilities of wheeled humanoid robots are gradually expanding from single-function delivery toward a combination of delivery, manipulation and human-robot interaction. In unmanned delivery, supermarket guidance, restaurant service and medical and elderly care scenarios, these robots must not only complete autonomous navigation but also possess stable grasping, placement and human-robot interaction capabilities. The dual-arm structure supports multi-item handling, two-handed delivery and more natural interactive motions, and emphasizes safety and reliability to improve service efficiency and user experience.
Wheeled humanoid robots in commercial service scenarios are evolving from perception-oriented to manipulation-oriented systems. For example, one service robot combines natural language interaction with anthropomorphic motions to cover greeting, cleaning and item delivery, demonstrating the application potential of wheeled humanoid robots in commercial services. Another company has used its self-developed wheeled humanoid robot to build an embodied intelligence retail solution for routine consumer operations, achieving fully autonomous operation across product recommendation, ordering and payment, precise grasping and in-person handover, driving a breakthrough application in retail services. Additional developers have released commercial-class wheeled dual-arm embodied intelligence service robots that integrate mobility, manipulation and interaction technology stacks with generalized manipulation capability. The release of these robots marks the beginning of commercial deployment for wheeled dual-arm service robots, and commercial service robots are entering a new era of embodied intelligence, which will drive the industry to upgrade from single-point functional substitution to full-scenario service coverage.
4.3 Household applications
In household applications, the potential uses of wheeled humanoid robots include item organization, living assistance and elderly care support, such as tidying, automatic cleaning and tool use. These tasks are typically highly unstructured and place higher demands on environmental understanding, manipulation stability and safety, and dual-arm collaboration offers clear advantages in handling diverse objects and complex operations.
In education, entertainment and emotional companionship, dual-arm mobile robots can enhance user experience through richer motion expression and interaction modalities. Combined with the needs of aging-friendly services, they have important application prospects in elderly companionship, daily assistance and safety monitoring. For example, one wheeled humanoid robot can clean, organize items, accurately locate objects and collect packages in household settings. Another wheeled humanoid robot can operate home appliances, store items and perform high-difficulty operations such as brewing tea and cooking. A further platform can provide multiple hardware and software adaptation options according to different customer scenarios, offering both emotional companionship and delicate operations such as bartending and dishwashing. Tencent Robotics X’s elderly-care robot supports safe physical human-robot interaction, can assist in caring for patients and the elderly, and can perform functions such as lifting and transferring, active suspension and folding and unfolding. The household scenario is the typical representative of unstructured environments and one of the ultimate markets for wheeled humanoid robots moving from industrial and commercial scenarios to civil applications. Although household applications still face challenges in cost, reliability and long-term autonomous operation, as perception, planning and embodied intelligence technologies continue to mature, wheeled humanoid robots are expected to play an increasingly important role in future home services.
5. Challenges and Bottlenecks
Despite notable progress in body design, mobile manipulation cooperative control and embodied intelligence methods, stable operation and engineering application in complex open environments still face multiple constraints. These issues are not concentrated in a single technical module but run through hardware integration, environmental perception, cooperative planning, task learning and safety standards. The review groups the main challenges into component and system integration, perception system robustness, coordinated dynamic planning of mobility and manipulation, generalization capability for long-horizon complex tasks, and safety and standards.
5.1 Core components and system integration
The overall performance ceiling of a wheeled humanoid robot is essentially determined by component integration level and system architecture design capability. The field faces a contradiction in component selection and system integration: high-performance, high-reliability actuation and perception units are needed, yet cost, mass, volume and endurance impose hard limits.
On energy consumption and power systems, the superposition of chassis motion, dual-arm operation, multimodal sensing and computing load makes peak and average power consumption significantly higher than that of single-arm or single mobile platforms, creating endurance shortages and thermal management difficulties. Energy consumption is mainly concentrated in high-torque servo drives, frequent start-stop and acceleration-deceleration motion, end-effector grasping or tool execution, and high-computing-power perception and planning units. Excessive motor margins, heavy mechanical structures and inefficient motion planning further amplify energy losses. Optimization therefore cannot remain at the component level but must simultaneously consider high-efficiency motors and drives, lightweight structural materials, regenerative braking and energy management, and coordinated path and motion optimization. For heterogeneous wheeled humanoid robots, a reasonably accurate energy prediction and management model is also needed, otherwise task interruption and scheduling failure can easily occur in multi-robot collaboration scenarios. Existing research still suffers from a disconnect between component-level optimization and system-level coordination, with insufficient attention to the dynamic variation of overall energy consumption under coupled mobility and manipulation. Future work should build an energy coordination optimization framework across component, system and task levels to alleviate the endurance bottleneck.
On integration constraints of actuation and computing units, supporting high-DOF dual-arm operation and whole-body cooperative control requires high-bandwidth servo drivers, torque and current sensing units, and high-performance graphics processors and AI accelerators. However, such components often come with large volume, high heat dissipation and power supply requirements and complex wiring, conflicting with the limited installation space of the mobile chassis and the lightweight requirements of the whole machine. Actuation units must balance dynamic response, impact resistance and long-term operational stability, while computing units must support perception, planning and control tasks simultaneously under limited power consumption, making coordinated integration among drives, sensing, computing power and power supply a key engineering difficulty. Existing solutions often sacrifice some performance in exchange for miniaturization and deployability, and still face significant engineering obstacles in packaging processes, thermal design and system adaptation. Future progress will depend on power semiconductors, heterogeneously integrated chips and high-density packaging technologies to promote efficient coordinated integration of actuation, sensing and computing units.
On end-effector dexterous actuators and multimodal sensing integration, dexterous hands and multifunctional end tools can significantly improve grasping and manipulation capability but introduce new complexity at the system integration level. Current high-DOF dexterous hands generally suffer from complex structures, dense actuation and sensing units, insufficient reliability and difficult maintenance, and in industrial and service scenarios, sensor drift, component wear and environmental interference can quickly degrade performance. Increasing degrees of freedom and multimodal perception usually means more actuation and data links, further raising wiring complexity, signal interference risk and control computation burden. For wheeled humanoid robots, the limited wrist space must simultaneously accommodate quick-change end tool interfaces, power and signal channels and necessary safety sensing units, making standardization and modularization of mechanical, electrical and software interfaces a key integration issue.
At the same time, to improve contact manipulation capability and human-robot collaboration safety, wheeled humanoid robots are gradually incorporating electronic skin, flexible tactile arrays, deformation sensors, distributed force sensors and multimodal perception networks. However, such sensors still have clear deficiencies in packaging, mechanical fatigue resistance, noise and drift control, and calibration compensation, and their integration compatibility with rigid structures and conventional control hardware is limited. Flexible tactile arrays based on conductive polymers and liquid metals are prone to conductive network damage and signal drift under large-deformation cyclic loading, and precise decoupling of pressure and shear force remains to be achieved. Distributed force sensing schemes such as fiber Bragg gratings and flexible capacitive arrays face engineering problems including bending fatigue, local failure and fault-tolerant reconstruction. In addition, multimodal sensor fusion involves real-time processing of high-dimensional data, communication bandwidth and latency control, requiring edge computing, hierarchical fusion architectures and real-time bus design to achieve efficient coordination between the sensing system and control hardware. The development of end effectors and multimodal sensing systems should therefore not simply pursue higher degrees of freedom and sensing density, but place greater emphasis on reliability, maintainability and integration cost in real scenarios.
At the whole-machine level, the system integration difficulty is not limited to whether components fit. High-DOF joints, multi-source perception systems and complex communication links increase capability but also significantly increase potential failure points. Under continuous operation, component wear, heat accumulation, loose connections, signal drift and software anomalies can all degrade system performance or even cause task failure. Compared with short-term validation in the laboratory, engineering applications pay more attention to stability and maintainability under repeated loads, continuous vibration and complex interference. Moreover, current wheeled humanoid robots still lack a well-developed standardization foundation in interface design, module replacement and system reuse. Different platforms differ greatly in drive units, sensor interfaces, end effector mounting methods and software communication protocols, increasing system development and maintenance costs and limiting platform expansion and industrial collaboration. Overall, the system integration difficulty of wheeled humanoid robots is not a matter of insufficient performance in a single component, but the lack of system-level design methods and standardization systems oriented toward the integrated characteristics of mobility and manipulation. This is the first key link that must be broken through as the technology moves from laboratory prototypes to large-scale engineering deployment.
5.2 Perception system robustness
Perception system robustness is an important foundation for autonomous operation in open environments and a prominent difficulty that distinguishes wheeled humanoid robots from traditional fixed manipulators and single mobile platforms. These robots must not only complete environmental perception and autonomous localization, but also continuously acquire multi-source information about the robot body, target objects, obstacles and surrounding personnel while mobility, manipulation and interaction occur in parallel. In unstructured scenes, the movement of people and changes in object poses can easily invalidate a pre-built map and cause localization loss. Improving the robustness of autonomous localization and mapping systems under drastic environmental change is therefore key to ensuring continuous execution of mobility and manipulation tasks. The current mainstream approach combines semantic information for dynamic feature removal with appearance-based fast relocalization to improve system recovery capability.
Beyond scene-level perception, wheeled humanoid robots also need finer state information during manipulation. When grasping transparent, reflective, flexible or stacked objects, single-source visual information often cannot provide a stable and reliable basis for operation. In dual-arm cooperative assembly, constrained manipulation and tool use, external vision alone also makes it difficult to accurately determine contact states and fit relationships. Cooperative fusion of vision, laser, inertial, force and tactile multimodal information is therefore a necessary path to improving perception robustness. It should be noted that multimodal perception is not equivalent to simply stacking sensors. Different sensors differ clearly in sampling frequency, noise characteristics, spatiotemporal synchronization and coordinate representation, and cross-modal fusion introduces new uncertainty propagation problems. Future research should address dynamic environment adaptation, long-term operational stability and multimodal information consistency to provide more reliable state input for subsequent planning and control.
5.3 Coordinated dynamic planning of mobility and manipulation
Strong dynamic coupling between the chassis and the two arms is the essential difficulty of coordinated planning. Unlike the traditional separated control mode of navigation first and manipulation second, wheeled humanoid robots usually need to adjust chassis pose, dual-arm trajectories and end-effector manipulation strategy simultaneously during motion to ensure continuity and efficiency of task execution. Inertial disturbances from chassis motion directly affect dual-arm manipulation precision, while reaction forces from dual-arm motion disrupt chassis force balance, and traditional separated control modes struggle to meet the needs of efficient operation. Coordinated planning currently faces three core contradictions: modeling precision versus computational efficiency in globally coupled models, trajectory optimality versus dynamic adaptability, and large-range operation versus whole-machine dynamic balance. Although existing whole-body model predictive control methods can partially solve coupled constraint problems, computing bottlenecks and insufficient robustness remain in highly dynamic scenarios. Achieving efficient solving of coupled models and real-time trajectory optimization while ensuring system stability is a core obstacle restricting expansion from structured scenarios to unstructured ones.
5.4 Generalization for long-horizon complex tasks
Insufficient generalization of complex tasks is the core bottleneck restricting wheeled humanoid robots from standardized scenarios to open general scenarios, spanning the full chain from task semantic understanding to skill transfer learning and physical execution. At the level of autonomous decomposition of long-horizon complex tasks, the core difficulty lies in semantic understanding of abstract instructions and causal logical reasoning. Existing large-model-based task planning methods can decompose fixed tasks in known scenarios, but in open scenarios with drastic changes in environment layout, unknown object attributes and ambiguous instructions, they struggle to stably generate sub-goal sequences that conform to physical rules and precondition logic. The essence is that existing models lack deep implicit understanding of causal relationships in the physical world and cannot achieve cross-scenario generalization of task logic.
At the level of cross-scenario skill transfer, the core constraints come from two dimensions. First, high-quality embodied datasets are systematically lacking: existing data collection is costly, annotation specifications are not unified, and most data focuses on single successful demonstrations without covering disturbances, failures and recovery processes, so models cannot obtain sufficient generalizable training samples. Second, there is the inherent gap between simulation and reality: existing simulation platforms cannot fully replicate the physical interaction details of the real world, causing policies trained in simulation to degrade sharply in real scenarios, especially in contact-intensive tasks such as fine manipulation and flexible object interaction, where the effect of sim-to-real transfer still has significant room for improvement. Improving long-horizon task generalization therefore requires not only algorithmic innovation but also a standardized dataset system, high-fidelity simulation environments and a closed-loop iterative learning framework of simulation pretraining, real-world fine-tuning and data feedback.
5.5 Safety and standards
The safety assurance system of wheeled humanoid robots is key to their large-scale application, and currently faces multiple challenges including physical interaction safety, data privacy protection and the lack of ethical norms. In physical safety, collision risk prevention and control in human-robot coexistence scenarios is the core issue: robots need intrinsic safety design and real-time obstacle avoidance capability. Although existing compliant actuation and force control algorithms can reduce harm, in highly dynamic scenarios the system still has deficiencies in collision identification, emergency stop and safety strategy switching, and unified and verifiable safety evaluation metrics are still lacking. In data privacy, sensitive information such as user images and voice collected by robots carries leakage risk, and dedicated encryption schemes for multimodal data remain immature. In ethical safety, the allocation of decision-making authority in emergency tasks and the division of responsibility for operational errors lack clear guidelines and legal norms, leaving room for dispute.
The lack of standards and industry specifications is another key bottleneck restricting large-scale application. Insufficient standardization of core components and software interfaces, including arm joint interfaces, sensor transmission protocols and API interfaces, increases integration costs. The absence of performance testing standards means there is no unified test and evaluation method for metrics such as manipulation precision and payload capacity, making it difficult to objectively reflect product performance. International standards for industrial robots are still difficult to fully cover systems that combine mobility and manipulation, and dedicated standards are still being developed. Domestic standards also mostly focus on industrial scenarios, while standards for commercial services and household scenarios lag relatively behind. There is an urgent need to establish a unified standards system covering multiple scenarios and the full product lifecycle to support standardized development and engineering deployment of wheeled humanoid robots.
6. Future Development Trends
The engineering deployment of wheeled humanoid robots faces a core contradiction between improving general manipulation capability and the reliability constraints of specific scenarios. In response, frontier breakthroughs are no longer limited to single-point iteration of hardware performance or algorithm effectiveness, but are forming a full-chain evolution direction spanning underlying data and simulation systems, core intelligence and manipulation capability, novel sensing hardware, and safety and human-robot symbiosis paradigms.
6.1 Multimodal data collection, processing, training and simulation
In the field of embodied intelligence, processing massive data and achieving efficient learning has become a key technical bottleneck, and the frontier trend is increasingly focused on simulation-first, data-driven and virtual-real closed loops. The core path is to use high-fidelity physical simulation platforms to build large-scale parallel virtual training grounds, and through domain randomization and domain adaptation techniques, inject realistic visual and dynamic diversity into synthetic data so that policies learn robustness to the complexity and uncertainty of the real world in advance. This process not only greatly reduces dependence on high-risk, high-cost trial and error in the physical world, but also achieves order-of-magnitude improvements in training efficiency through distributed reinforcement learning architectures.
However, relying solely on simulation cannot fully bridge the sim-to-real gap. Frontier research is therefore committed to building automated, closed-loop data pipelines: policies pretrained in simulation are deployed on real robots for silent trials or human-supervised interaction, and the collected real-world trajectories are automatically cleaned and annotated to form high-value datasets. These data can be used directly for online imitation learning or offline reinforcement learning to fine-tune policies, and can also be used in a simulation-to-real-to-simulation loop to correct the physical parameters of the simulation model, making the virtual environment continuously approach the real world.
The ultimate form of this trend is a robot digital twin system that is synchronized with the physical entity in real time and possesses predictive capability. It makes it possible to safely and rapidly test new skills, plan complex task chains and rehearse anomaly recovery schemes in virtual space, and then seamlessly issue verified plans to the physical entity for execution. By exploring boldly in simulation and executing robustly in reality, the development, testing and deployment cycle of robot skills is greatly compressed, enabling robot learning to shift comprehensively from the traditional mode of manual programming and limited experimentation to a scalable new paradigm driven jointly by large-scale synthetic data and continuous real-world feedback.
6.2 General intelligence and highly humanoid dexterous manipulation
General intelligence and highly humanoid dexterous manipulation are core development directions for wheeled humanoid robots, accelerating from traditional dedicated skill-driven approaches toward large-model-enabled general embodied intelligence. Early robot dexterous manipulation relied on task-specific models with limited generalization capability. General embodied intelligence based on large language models fuses multimodal perception information such as vision, touch and force to achieve end-to-end mapping from natural language instructions to manipulation tasks, improving the naturalness of human-robot interaction and task adaptation flexibility. Combined with reinforcement learning and sub-skill controller co-training paradigms, robots can autonomously acquire human-like manipulation behaviors such as finger gaiting and multi-finger coordination. The core is to optimize exploration efficiency in complex state spaces through value-guided exploration strategies, allowing high-precision dexterous manipulation capability to emerge without manual demonstration. At the hardware level, structural optimization and sensor integration of multi-DOF dexterous hands provide the physical foundation, while domain randomization in simulation environments narrows the sim-to-real gap, enabling trained policies to be deployed stably on physical robots. This co-evolution path drives dexterous manipulation from task-specific toward general intelligence with environmental self-adaptation and instruction self-understanding, laying a key foundation for large-scale application of wheeled humanoid robots in complex scenarios.
6.3 Novel multidisciplinary sensors
With the development of materials science and neural engineering, multidisciplinary novel sensor fusion is giving robots more advanced perception capabilities. In addition to conventional force and vision sensing, electronic skin based on flexible electronics can give robots large-area tactile perception, enabling safe collision detection and fine manipulation in human-robot interaction. In addition, the introduction of electromyographic interfaces and brain-machine interfaces allows robots to directly interpret human motion intentions. This fusion of biological-signal-based intention recognition and novel sensors will greatly expand the application boundaries of dual-arm humanoid robots in medical rehabilitation and human-robot collaboration.
Under the trend of simultaneous evolution at both the perception and expression ends, facial expression simulation is also gradually becoming an important extension of multimodal sensing and interaction. The development of flexible, stretchable sensing and actuation materials allows robot faces to integrate high-density facial electronic skin that can sense touch, pressure and temperature while presenting delicate expression changes through deformable structures. These facial expression signals, jointly fused with vision, voice, touch and bioelectric signals, can be incorporated into multimodal perception-decision frameworks for emotion recognition and emotional expression, thereby enhancing the robot’s social presence and acceptability for long-term coexistence in companionship, education and rehabilitation scenarios.
Through deep fusion of multi-source novel sensors including electronic skin, bioelectric interfaces and facial expression simulation, wheeled humanoid robots are gradually moving from physical collaboration tools toward intelligent partners with situational understanding and emotional interaction capability, which will greatly expand their application boundaries in medical rehabilitation, human-robot collaborative manufacturing and daily services.
6.4 Safety, ethics and human-robot-environment symbiosis
Human-robot-environment symbiosis and the standardization of safety and ethics are the core support for wheeled humanoid robots to embody people-centered and human-centric intelligent manufacturing concepts and to build a new paradigm of human-machine integration. As application scenarios extend to all sectors of society, achieving harmonious coexistence among humans, robots and the environment has become a key proposition for technological iteration and industrial development. To achieve this, safety design and ethical norm construction are being advanced comprehensively from technical, institutional and social dimensions, gradually forming a development system suited to the coordinated mobility and manipulation characteristics of these robots.
At the technical level, safety optimization focuses on risk prevention and control in both mobility and manipulation scenarios. The deep combination of multimodal sensor fusion and intelligent decision algorithms allows robots to accurately perceive potential risks in dynamic environments and avoid safety hazards such as collisions and overloads by adjusting motion trajectories and manipulation force in real time. Intrinsic safety design for mobility-manipulation coupling continues to upgrade, and the application of compliant actuation technology and adaptive control strategies gives robots better compliance and fault tolerance in human-robot interaction, reducing physical interaction risk from the underlying architecture and laying a technical foundation for human-centric intelligent manufacturing.
At the ethical level, as wheeled humanoid robots participate in more complex tasks, issues such as decision-making boundaries, responsibility allocation and data privacy protection have attracted widespread attention. Governments, international organizations and academia are jointly promoting targeted regulation, establishing ethical guidelines and legal frameworks adapted to the technical characteristics of these robots around core topics such as behavioral constraints in mobile operations, data collection and use specifications, and fault responsibility determination, ensuring that robot behavior always conforms to social ethics and public interest.
At the social level, deepening human-machine integration depends not only on technical and institutional safeguards but also on upgrading society’s understanding of wheeled humanoid robots. Through science communication and application practice, public trust in robots can be gradually established, while the role of robots in the social division of labor should be clearly defined to ensure controllability and transparency of their behavior, achieving synchronous improvement of technological progress and social acceptance. As safety and ethical systems continue to improve, wheeled humanoid robots will integrate more smoothly into human production and life.
7. Industrial Development Considerations
As the core engineering carrier for the deployment of embodied intelligence technology, the wheeled humanoid robot industry is deeply coupled with frontier technological evolution and currently stands in a critical window between technical validation and practical application. A preliminary industrial form covering core components, system integration and application validation has taken shape, but it still faces practical problems such as poor transformation of core technologies, insufficient depth of scenario deployment and an incomplete industrial support system. Development paths must therefore be built around three core dimensions: technological innovation, scenario deployment and ecosystem construction.
7.1 Strengthening collaborative technological innovation
Focusing on the strongly coupled technical characteristics of mobility and manipulation, priority should be given to unified dynamic modeling, whole-body motion planning for highly redundant systems and coordinated control in dynamic environments. In parallel, key hardware such as high-power-density miniaturized servo drives, high-density flexible tactile sensing and highly reliable dexterous hand integration should be tackled, and overall energy management and lightweight design should be optimized to resolve the core contradiction among performance, volume and energy consumption. A collaborative innovation system based on deep industry-academia-research integration should combine the basic research accumulation of universities and research institutions with the engineering and scenario validation advantages of leading enterprises, build dedicated innovation platforms and open the transformation chain from laboratory results to industrialization. Building public infrastructure such as open-source simulation training platforms and standardized embodied manipulation datasets, and unifying data collection and annotation specifications, will lower the innovation threshold for the industry and drive the transition from single-point technological breakthroughs to system-level capability improvement.
7.2 Deepening scenario penetration and innovating service models
The technical and industrial value of wheeled humanoid robots must ultimately be realized through scenario deployment, following the principles of gradient penetration and demand orientation to build a stepped scenario expansion path. Industrial manufacturing should be cultivated first, targeting new energy vehicles, consumer electronics, logistics sorting, patrol and guarding, and inspection, focusing on clear needs such as collaborative assembly, material handling and equipment inspection, and building standardized solutions adapted to production line rhythms to secure a basic revenue foundation.
At the same time, commercial service scenarios such as retail, hospitality, medical rehabilitation and elderly and disability care should be steadily expanded, optimizing human-robot interaction safety and multi-task adaptability, and creating replicable benchmark demonstration projects. Household service scenarios should be laid out prospectively, driving high-reliability product iteration. A product development model of general body platform plus modular configuration plus scenario-specific process packages can balance generalization and customization requirements, while new business models such as robotics-as-a-service and integrated leasing and maintenance should be promoted to build a commercial path linking hardware, software, services and data.
7.3 Improving the industrial ecosystem
A sound industrial ecosystem is the core support for large-scale, high-quality development, and should be built around three core directions: standards, supply chain collaboration and talent supply. An industry standards system covering the full product lifecycle should be established quickly, unifying hardware interfaces, communication protocols and software API specifications, improving product performance testing methods and evaluation systems, and formulating dedicated standards for physical safety and data privacy in human-robot coexistence scenarios, while actively participating in international standard setting to enhance international influence in this field. Collaboration along the entire industrial chain should be deepened, cultivating a pattern led by leading enterprises with specialized and sophisticated supporting suppliers, strengthening domestic supply capacity for core components, ensuring supply chain autonomy and controllability and improving overall industrial chain resilience. A multi-level interdisciplinary talent development system should be built through interdisciplinary programs, university-enterprise joint training and vocational skills training, creating a talent pipeline covering research, engineering and application to provide stable intellectual support for sustained industrial development.
8. Conclusions
Research on wheeled humanoid robots has formed a relatively complete technological and industrial development framework. Its core dimensions include body hardware structure and software architecture, mobile chassis environmental perception and motion planning, task understanding and dual-arm dexterous manipulation, data acquisition and simulation training for complex scenarios, and embodied intelligence control paradigms with sim-to-real transfer. The central research focus is on unified dynamic modeling of the wheeled chassis and dual-arm system, cooperative motion planning for highly redundant systems, multimodal sensor information fusion and compliant control. Technological evolution shows a clear trend from separately controlled subsystems toward integrated whole-body control of mobility and manipulation, and the combined application of multi-sensor fusion, high-fidelity simulation and intelligent algorithms has become the core path for breaking through existing technical bottlenecks and provides important equipment support for the physical deployment of embodied intelligence.
Wheeled humanoid robots have already completed engineering pilot applications in industrial manufacturing, commercial services and special operations, and a development pattern in which technology research and engineering deployment advance in parallel has taken shape, with a preliminary industrial system covering core components, complete machine development and scenario applications. Large-scale engineering application is still constrained by the integration of high-performance core components, coordinated planning in dynamic environments, dataset collection and training, task generalization in unstructured scenarios, and an incomplete industry standards system. As these technologies continue to advance and the industrial ecosystem matures, wheeled humanoid robots are expected to achieve broader large-scale application in intelligent manufacturing and service fields, carrying significant academic research value and engineering application prospects for the ongoing development of embodied intelligence.
