Embodied Intelligence Beyond Limbs: The Quiet Contest Over Robot Brains at the 2026 World Robot Conference

At the 2026 World Robot Conference in Beijing E-Town, more than 300 companies and more than 3,000 exhibits gathered in one place. On the exhibition floor, robots worked. On the conference floor, practitioners argued about the same question: what should a robot’s brain look like? The question is central to embodied intelligence, because the arms and legs of humanoid robots have already been pushed to an extreme, and the industry’s attention is now moving upward.

If last year the industry was still debating whether end-to-end learning was the answer, this year the more urgent question became another one: VLA, or vision-language-action models, world models, and unified models, which is the ultimate form of the robot brain? For embodied intelligence, the debate is no longer only about architecture. It is about data, physical understanding, evaluation, and the last mile from brain to body.

1. The Heat Cools, but the Question Sharpens

The most visible change at the 2026 World Robot Conference was that the debate over technical routes became noticeably quieter. If one word were used to summarize the technical atmosphere, it might be cooling. Last year, VLA signs were everywhere on booths and forums. This year, they were visibly fewer. End-to-end learning was no longer treated as a universal answer. The new concept that took over was the world model, but its heat did not mean agreement. Different companies defined it differently: generating the future, predicting causality, understanding through events, or simulating in latent space.

For embodied intelligence, the cooling is not a retreat. It is a form of discipline. The industry is moving from route disputes to the data question. The strengths and weaknesses of an algorithm architecture must ultimately be tested by data. Rather than arguing VLA versus world model, the more basic questions are: where does data come from, and how is it collected?

Su Hang, an associate researcher in the Department of Computer Science and Technology at Tsinghua University, gave a widely cited judgment during a roundtable at the World Robot Conference. In his view, VLA and world models are not an either-or relationship. For data with relatively high quality and clear action and task labels, VLA can directly establish a mapping among perception, language, and action. When facing larger-scale video data from broader sources, world models can help the model learn the relationship between environmental changes and behavioral results.

Su Hang predicted that the future is more likely to bring fusion architectures. At different time scales, such architectures would combine VLA’s action generation capability with a world model’s state prediction and feedback correction capability. For embodied intelligence, this is an important shift. The debate is no longer about replacing one route with another. It is about combining strengths in a system that can perceive, reason, act, and correct itself.

From the practice of several companies, VLA has not been eliminated. It has changed from an independent technical route into a system foundation. Qinglang Intelligent released its new-generation model KOM3.0 at the World Robot Conference. It was presented as the world’s first service-industry VLA architecture integrating a latent-space world model. Its core breakthrough is that the robot internally rehearses physical results before executing an action. It predicts the movement trajectory of an object after force is applied and the causal chain of multi-step operations. The goal is to avoid failure at the source rather than relying on traditional trial-and-error execution.

The logic is straightforward for embodied intelligence. VLA solves the problem of being able to see and being able to work. The world model solves the problem of predicting consequences and reducing trial and error. When the two are stacked, the robot can think once before it works. This is a shift from reactive control toward predictive control, and it is one of the clearest expressions of embodied intelligence in the service robotics field.

Gao Jiyang, CEO of Xinghaitu, offered a more radical judgment. He said VLA and the world model have the same core. Both are based on Transformer architecture and self-supervised pretraining. The difference is only in the method of self-supervision. In his view, the two are homologous and symbiotic, and they will move toward fusion in the future. This view challenges the idea that the two routes are separate camps. For embodied intelligence, it suggests that architectural boundaries may matter less than the data and objectives that drive learning.

Lexiang Technology’s Yuandian Robot stressed that a single model can hardly solve all the problems of embodied intelligence in the real physical world on its own. VLA solves the mapping from vision and language to action, but in unstructured scenes such as homes, it still needs a data platform, a training platform, and a body system to collaborate. Together, these elements can build a capability loop that continuously evolves. This is a reminder that the robot brain is not only a model. It is a system.

Everyone at the conference was talking about world models, but they were not necessarily talking about the same model. Huang Yuanhao, founder of Orbbec, said at the main forum that large language models can only solve poetry and distance, while world models are what let robots truly work. In his view, GPT made human-machine dialogue a reality, but the world model must solve the labor problem of 8 billion people and liberate humans from repetitive labor. This framing places embodied intelligence at the center of economic and social change, not merely at the center of a technical benchmark.

Hu Luhui, founder of Zhicheng AI, pointed to a more fundamental issue. In his view, the world model must solve the most core proposition of artificial intelligence: understanding the physical world. If multimodal large models remain at the perception level, the world model must enter the deep water of causal reasoning and decision-making. Hu broke understanding the physical world into three levels in an interview: cognition of the physical properties of objects, cognition of the dynamic change laws of the environment, and task-oriented behavior planning ability.

In his view, the world model is indispensable underlying infrastructure for humanoid robots. A robot observes its surrounding environment and outputs its next action. It must rely on physical rules and dynamic change laws, and it must rely on causal reasoning to make judgments. For embodied intelligence, this means intelligence cannot be only linguistic or visual. It must be physical.

On the technical route, Zhicheng AI chose the JEPA direction, which stands for Joint Embedding Predictive Architecture. It learns physical laws in abstract representations. The company’s self-developed Chengling Physical World Model can learn latent representations of physical world behavior from raw inputs such as vision and proprioception. This allows the robot to complete a kind of mental simulation before executing a task. In embodied intelligence, this is an attempt to act internally before acting externally.

Wujie Dynamics released MWA, described as the world’s first long-horizon bidirectional physical causal chain latent-space world model. It completes unified representation and a reasoning closed loop of multimodal information in latent space. Unitree Technology focused on the video generation path. Its WALL-WM model aligns vision, language, and action through events, enabling the robot to simultaneously understand the process, state changes, and results.

Jia Kui, founder of Cross-dimensional Intelligence, used even more ambitious language. He described the world model as the pearl on the crown of generative AI and the endgame of AI unification. In his view, it can not only drive robots to work, but also generate content with absolutely correct 3D physical geometry. These statements show how broad the world model concept has become, and they also show why the industry still lacks a shared definition.

Industry insiders also pointed to real challenges. Wang Xiaogang, chairman of Daxiao Robot, admitted during the World Robot Conference that the world model has gone through three stages of evolution: from data augmentation to simulation, and then to edge deployment. But it is currently still climbing. The leap from static scenes to dynamic real environments is far harder than expected. The bigger challenge is that the industry has not unified the underlying definition of the world model. Wang said that what everyone is doing is still very rough in terms of underlying alignment, and even evaluation standards are not fully unified. For embodied intelligence, the lack of shared evaluation is a bottleneck.

2. From Patchwork to Collaboration: Unified Models and New Routes

Beyond the debate between VLA and world models, another route is surfacing: the unified model. Beijing Humanoid Robot Innovation Center launched Pelican-Unify at the World Robot Conference. It was described as the nation’s first unified representation embodied world model. Its core breakthrough is that it no longer splits visual understanding, action manipulation, and world prediction into independent modules. Instead, the three share the same representation system and co-evolve.

According to the center’s explanation, whether the input is text, image, video, or action code, it is aligned to a shared representation space. Reasoning and planning are completed within the same framework, and actions are ultimately generated end-to-end. This makes understanding, reasoning, planning, and action no longer a serial patchwork, but an integrated closed loop. For embodied intelligence, this is an attempt to overcome fragmentation at the architectural level.

Xiong Youjun, CEO of Beijing Humanoid Robot Innovation Center, said in an interview that traditional VLM models can see and think but cannot act. VLA models can act but cannot rehearse. World models can rehearse but are not good at long-horizon reasoning. Pelican-Unify unifies vision, text, and action with the same paradigm. In May 2026, it topped the WorldArena global authoritative benchmark. This means embodied intelligence is moving from functional patchwork to collaborative evolution, he suggested.

In addition to the VLA plus world model fusion route and the unified model route, some companies chose entirely different directions. Zhipingfang launched and open-sourced its original brain-inspired architecture embodied model NeuroVLA at the World Robot Conference. It introduces the coordination mechanism of the human brain’s cortex, cerebellum, and spinal cord into the robot control system. The cortex is responsible for semantic understanding and task planning. The cerebellum is responsible for high-frequency motion coordination and dynamic correction. The spinal cord is responsible for millisecond-level action execution and safety reflexes.

Zhang Peng, co-founder of Zhipingfang, explained the logic in an interview. He said VLA and world models are not a matter of path selection, but a matter of what mode to choose across the entire technical path to enhance systematic capability. He said the company started doing end-to-end learning in 2023, introduced a fast-slow dual system in 2024, and launched NeuroVLA in 2026. The core is to make the system more and more like a human in thinking and acting. For embodied intelligence, this route stresses biological plausibility and layered control.

Shenyi Robotics proposed the concept of self-evolution. Huang Kangyao, co-founder of the company, said at a World Robot Conference roundtable that embodied scenarios are in real physical space. Unlike large language models, where text is traversable, real space also requires robots to output and generate continuous action control. This means data cannot be fully collected. Shenyi internally proposed two dimensions: L4E, which stands for Learning for Embodiment, meaning learn first and then work, and E4L, which stands for Embodiment for Learning, meaning the robot continuously learns in real work.

On the model route, Shenyi does not follow VLA or the world model. It uses an original architecture internally called a hybrid physical expert model. The company attempts to integrate physical law understanding, motion prediction, obstacle avoidance, 3D estimation, and scene understanding into the same self-evolution paradigm. This reflects another frontier of embodied intelligence: continuous learning in the physical world.

3. The Strongest Signal: Routes Not Converged, Direction Clearer

The observation from the 2026 World Robot Conference is that the strongest signal is not a single winning route. Technical routes have not yet converged, but the direction is becoming clearer. Industry consensus is shifting from the question of which is the ultimate route for the robot brain to more pragmatic questions. Whether a company takes the VLA plus world model fusion route, the unified model collaboration route, or the brain-inspired bionic route, it ultimately returns to the same starting point: data.

Embodied intelligence is evolving from moving on stage and running on tracks to being used in homes and working in factories. In this evolution, the choice of technical route is important, but more critical is who can run data through in real scenes, run the closed loop smoothly, and truly walk the last mile from brain to body.

To organize the landscape, the following table summarizes the major routes discussed at the conference. It does not rank them. It shows how each route addresses different parts of embodied intelligence.

Table 1. Technology routes discussed for embodied intelligence at the 2026 World Robot Conference
Route Core proposition Representative examples mentioned Open question
VLA plus world model fusion VLA maps perception, language, and action; world model predicts states and consequences; combine them at different time scales Qinglang KOM3.0; Xinghaitu perspective; Su Hang’s fusion view How to align data and evaluation across modules
Unified representation model Vision, text, and action share one representation system and evolve together Beijing Humanoid Robot Innovation Center Pelican-Unify Can unified representation support long-horizon reasoning and real-world action
World model as foundation Understand the physical world, reason causally, and simulate mentally before action Zhicheng AI JEPA and Chengling; Wujie Dynamics MWA; Unitree WALL-WM; Cross-dimensional Intelligence Definition and evaluation standards remain unsettled
Brain-inspired control Cortex, cerebellum, and spinal cord coordination for planning, correction, and execution Zhipingfang NeuroVLA How to scale biological principles into general embodied intelligence
Self-evolution and hybrid physical experts Continuous learning in real physical work; integrate physical law, prediction, avoidance, 3D estimation, and scene understanding Shenyi Robotics L4E and E4L; hybrid physical expert model How to collect and use data that can never be fully collected

4. Route-by-Route: How Each Camp Defines the Robot Brain

The VLA camp has not disappeared. Instead, it has become more mature. VLA is now often treated as a foundation for embodied intelligence rather than a complete answer. Qinglang’s KOM3.0 shows how VLA can be combined with a latent-space world model to support service-industry tasks. The robot first rehearses physical results internally. It predicts how an object will move after force is applied. It predicts the causal chain of multi-step operations. This is a practical way to reduce trial and error in embodied intelligence.

Su Hang’s view gives the VLA route a clear role. When data quality is high and actions and tasks are clearly labeled, VLA can directly build a mapping between perception, language, and action. This is especially useful in controlled tasks where action labels are available. The limitation is that VLA alone may not predict long-horizon physical consequences well. That is where the world model enters.

The world model camp is broader and more diverse. Huang Yuanhao of Orbbec framed the world model as the key to making robots truly work. Hu Luhui of Zhicheng AI framed it as the path to understanding the physical world. Zhicheng AI’s JEPA direction learns physical laws in abstract representations. Wujie Dynamics’s MWA works in latent space and claims a long-horizon bidirectional physical causal chain. Unitree’s WALL-WM uses video generation and event alignment to connect vision, language, and action. Cross-dimensional Intelligence’s Jia Kui calls the world model the pearl on the crown of generative AI and the endgame of AI unification.

For embodied intelligence, the world model route has a powerful appeal. A robot cannot rely only on pattern recognition. It must understand physical rules, dynamic changes, and causal relationships. It must predict what will happen before it acts. However, the route also faces serious challenges. Wang Xiaogang of Daxiao Robot said the world model is still climbing from static scenes to dynamic real environments. The industry has not unified its underlying definition. Even evaluation standards are not fully unified. Without shared definitions, progress in embodied intelligence becomes difficult to compare.

The unified model camp tries to solve fragmentation. Pelican-Unify from Beijing Humanoid Robot Innovation Center does not separate visual understanding, action manipulation, and world prediction. It aligns text, image, video, and action code into a shared representation space. It then performs reasoning and planning in the same framework and generates actions end-to-end. Xiong Youjun’s summary is useful for embodied intelligence: VLM can see and think but cannot act; VLA can act but cannot rehearse; world model can rehearse but is not good at long-horizon reasoning. Pelican-Unify aims to close the loop. Its May 2026 top ranking on WorldArena is a signal that unified representation is competitive in embodied intelligence.

The brain-inspired camp adds a different layer. Zhipingfang’s NeuroVLA introduces cortex, cerebellum, and spinal cord coordination. The cortex handles semantic understanding and task planning. The cerebellum handles high-frequency motion coordination and dynamic correction. The spinal cord handles millisecond-level action execution and safety reflexes. Zhang Peng said the company began with end-to-end learning in 2023, introduced a fast-slow dual system in 2024, and launched NeuroVLA in 2026. The goal is to make the system more human-like in thinking and acting. For embodied intelligence, this route stresses layered control, safety, and real-time response.

The self-evolution camp focuses on a paradox. A robot must learn before it works, and it must also learn while it works. Shenyi Robotics calls these two dimensions L4E and E4L. L4E means Learning for Embodiment, or learning first and then working. E4L means Embodiment for Learning, or continuous learning in real work. Because real space requires continuous action control, data cannot be fully collected in advance. Shenyi’s hybrid physical expert model integrates physical law understanding, motion prediction, obstacle avoidance, 3D estimation, and scene understanding into a self-evolution paradigm. This is a practical response to data scarcity and non-stationarity in embodied intelligence.

5. Data Is the Common Starting Line for Embodied Intelligence

The data question is not narrow. For VLA, data must contain clear action and task labels. For world models, data may be larger-scale and from broader sources, such as video. For unified representation, data must be aligned across text, image, video, and action. For brain-inspired control, data must support high-frequency correction and millisecond-level execution. For self-evolution, data cannot be fully collected in advance because real space generates continuous action control.

This is why embodied intelligence cannot simply copy the recipe of large language models. Text on the internet is vast and traversable. Physical interaction is not. A robot in a home or factory faces changing objects, changing lighting, changing human behavior, and changing task requirements. The data needed for embodied intelligence is not only observational. It is interactive and outcome-based.

Su Hang’s distinction helps clarify the data question. When data quality is high and actions and tasks are clearly labeled, VLA can directly establish a mapping among perception, language, and action. When facing larger-scale and broader-source video data, the world model can help the model learn the relationship between environmental changes and behavioral outcomes. Therefore, the future may be a fusion architecture. At different time scales, VLA’s action generation can be combined with the world model’s state prediction and feedback correction. This is a data-centric view of embodied intelligence.

Qinglang’s KOM3.0 illustrates how predictive rehearsal can reduce trial and error. Before executing an action, the robot internally rehearses physical outcomes. It predicts object motion trajectory after force is applied and the causal chain of multi-step operations. This requires data that connects perception, action, and physical consequence. For embodied intelligence, such data is more than observation. It is interaction and result.

Huang Yuanhao of Orbbec framed the world model as a solution to the labor problem of 8 billion people. That scale implies data from many environments and tasks. The challenge is not only collecting data but organizing it, labeling it, and closing the loop. Lexiang Technology’s Yuandian Robot noted that a single model cannot solve all problems in the real physical world. A data platform, a training platform, and a body system must collaborate to build a continuously evolving capability loop. This is a system-level view of embodied intelligence.

Zhicheng AI’s three levels of understanding the physical world also imply different data needs. Cognition of object physical properties, cognition of environment dynamic laws, and task-oriented behavior planning require different data and representations. The JEPA direction learns physical laws in abstract representation. It uses vision and proprioception to form latent representations. This is an attempt to make embodied intelligence learn from raw inputs without relying only on labeled action data.

Wujie Dynamics’s MWA and Unitree’s WALL-WM show two paths: latent-space causal reasoning and video generation with event alignment. Both try to unify multimodal information. Cross-dimensional Intelligence’s Jia Kui described the world model as the pearl on the crown of generative AI and the endgame of AI unification. Such ambition raises the evaluation question. Wang Xiaogang of Daxiao Robot said the world model has moved from data augmentation to simulation to edge deployment but is still climbing. Definitions and evaluation standards are not fully unified. For embodied intelligence, without shared benchmarks, progress is hard to measure.

Beijing Humanoid Robot Innovation Center’s Pelican-Unify tries to solve fragmentation through shared representation. It aligns text, image, video, and action code into a shared space. It then reasons and plans in the same framework and generates actions end-to-end. Xiong Youjun’s description of VLM, VLA, and world model limitations shows why unified representation is attractive. VLM can see and think but not act. VLA can act but not rehearse. World model can rehearse but is not good at long-horizon reasoning. Pelican-Unify aims to close the loop. Its May 2026 top ranking on WorldArena is a signal that unified representation is competitive in embodied intelligence.

Zhipingfang’s NeuroVLA adds another layer. The cortex handles semantics and planning. The cerebellum handles high-frequency coordination and correction. The spinal cord handles millisecond execution and safety reflexes. This brain-inspired design reflects embodied intelligence’s need for both thought and reflex. Zhang Peng said the company began end-to-end learning in 2023, introduced a fast-slow dual system in 2024, and launched NeuroVLA in 2026. The goal is a system that is more human-like in thinking and acting.

Shenyi Robotics’s L4E and E4L capture a paradox. A robot must learn before work, and learn while working. Because real space requires continuous action control, data cannot be fully collected. Its hybrid physical expert model integrates physical law understanding, motion prediction, obstacle avoidance, 3D estimation, and scene understanding into a self-evolution paradigm. This is a practical response to data scarcity and non-stationarity in embodied intelligence.

6. Evaluation and the Missing Common Language

The World Robot Conference made another issue clear: embodied intelligence lacks a common language for evaluation. Wang Xiaogang of Daxiao Robot said the underlying definitions of world models are still not unified. Even evaluation standards are not fully unified. This is not a minor academic concern. In embodied intelligence, evaluation determines what counts as progress. If different teams define world models differently, they may not be measuring the same capability.

The world model has already gone through several stages, according to Wang Xiaogang. It has moved from data augmentation to simulation and then to edge deployment. But it is still climbing. The leap from static scenes to dynamic real environments is harder than expected. For embodied intelligence, this means a model that performs well in simulation may fail in a home, a factory, or a public space. Dynamic real environments introduce uncertainty, human behavior, changing objects, and safety constraints.

A shared evaluation framework for embodied intelligence would need to cover several dimensions. It would need to measure physical correctness, causal reasoning, long-horizon planning, real-time control, safety, and generalization across environments. It would need to compare VLA, world models, unified models, brain-inspired architectures, and self-evolution systems. Pelican-Unify’s top ranking on WorldArena in May 2026 is one benchmark signal, but the broader field still needs more common tests.

Evaluation is also connected to data. If data is collected differently, models are evaluated differently. If tasks are defined differently, results are not comparable. For embodied intelligence, the missing common language is not only a technical problem. It is an ecosystem problem. It affects investment, research priorities, and deployment decisions.

7. From Conference Floor to Factory Floor

The 2026 World Robot Conference showed that embodied intelligence is moving from demonstration to deployment. More than 300 companies and more than 3,000 exhibits gathered in Beijing E-Town. Robots worked on the exhibition floor. But the conference floor debate showed that the real challenge is not only making a robot move. It is making a robot understand, predict, decide, and act reliably in the physical world.

Embodied intelligence is evolving from moving on stage and running on tracks to being used in homes and working in factories. This evolution requires more than a single model. It requires data pipelines, training platforms, body systems, safety layers, and evaluation standards. It requires VLA for perception-language-action mapping. It requires world models for physical prediction and causal reasoning. It requires unified representations to reduce fragmentation. It requires brain-inspired control for fast correction and safety. It requires self-evolution for continuous learning in real work.

The last mile from brain to body is not a metaphor. It is a practical gap. A robot may understand a command, but still fail to grasp an object. It may predict a trajectory, but still fail to correct in time. It may rehearse an action internally, but still face an unexpected human movement. Embodied intelligence must close these gaps in real time. The company that can run data through in real scenes and run the closed loop smoothly will be the one that walks the last mile.

8. What to Watch Next

At the 2026 World Robot Conference, the debate did not produce a single winner. Instead, it produced a more practical agenda for embodied intelligence. The next milestones are likely to include the following:

  • Data pipelines that connect real-world interaction, simulation, and deployment.
  • Evaluation standards that compare VLA, world models, unified models, brain-inspired systems, and self-evolution architectures.
  • Fusion architectures that use VLA for action generation and world models for state prediction and feedback correction at different time scales.
  • Unified representation systems that align vision, text, action, and physical geometry.
  • Brain-inspired control that combines semantic planning, high-frequency correction, and millisecond execution.
  • Self-evolution methods such as L4E and E4L that allow robots to learn before work and during work.
  • Real-world closed loops in homes and factories, where embodied intelligence moves from demonstration to daily use.

The 2026 World Robot Conference made one point clear: embodied intelligence is no longer only about limbs. The arms and legs have been pushed to an extreme. Attention has moved upward to the brain. But the brain will not be defined by a single label. VLA, world models, unified models, brain-inspired architectures, and self-evolution models are competing and converging. The decisive question is not which label wins. The decisive question is which system can collect the right data, build the right representation, run the right closed loop, and deliver reliable action in the physical world.

That is the last mile from brain to body for embodied intelligence. It is also the next frontier for the entire robotics industry. The companies that can connect data, models, bodies, and real-world feedback will shape how embodied intelligence moves from the conference floor to the factory floor, and from the factory floor to everyday life.

Scroll to Top