Embodied Intelligence and the Race for the Robot Brain: VLA, World Models, Unified Architectures, and the Data Test

At the 2026 World Robot Conference, more than 300 companies and more than 3,000 exhibits gathered in Beijing Yizhuang. On the exhibition floor, robots were working. Off the floor, practitioners were arguing about a single question: what should the robot brain look like? If last year the central debate was whether end-to-end models were the answer, this year the question became sharper. Among VLA, or vision-language-action models, world models, and unified models, which represents the ultimate form of the robot brain? The search for that answer is now shaping the entire field of embodied intelligence.

The debate itself has changed character. The most visible shift at the 2026 World Robot Conference was not the arrival of a single winning architecture. It was the decline of loud, simplistic route claims. VLA signs that had appeared everywhere on stages and exhibition booths a year earlier were noticeably fewer. End-to-end was no longer treated as a universal answer. The new concept that took over much of the discussion, the world model, was undeniably popular, but different companies attached different meanings to it. Some described a system that generates the future. Some described one that predicts causality. Some described one that understands events. Others described one that reasons in latent space. The vocabulary was shared, but the technical targets were not.

This cooling of concept-level enthusiasm should not be mistaken for a loss of momentum. The field of embodied intelligence is not retreating. It is becoming more disciplined. The industry is moving from a contest over slogans to a deeper question about data. Which architecture is superior can only be tested by data. Before arguing endlessly about VLA versus world models, the more fundamental questions are simpler and harder: where does the data come from, and how is it collected? That question now sits at the center of embodied intelligence.

1. At WRC 2026, the Robot Brain Becomes the Main Battlefield

The 2026 World Robot Conference made one thing clear: the body of the robot is no longer the only frontier. Robotic limbs, joints, hands, and mobility systems have already been pushed to an extreme level of competition. The industry’s attention is moving upward, toward the brain. That brain is not a single chip or a single algorithm. In the context of embodied intelligence, the robot brain is a layered system that must connect perception, language, action, prediction, reasoning, and control.

The conference brought together more than 300 companies and more than 3,000 exhibits. Yet the most important battle was not visible in a single product demonstration. It was taking place in the assumptions behind those products. One camp argues that VLA models can become the foundation of embodied intelligence because they directly map visual input and language instructions into actions. Another camp argues that world models are indispensable because robots must predict the consequences of their actions before they act. A third camp argues that the real answer is a unified model that avoids stitching separate modules together. Each position reflects a different understanding of what embodied intelligence actually requires.

In previous years, the debate often reduced to a binary choice. Was end-to-end learning the correct path? Was a modular pipeline more reliable? In 2026, the question is no longer so narrow. The real issue is not whether one model family can replace all others. The real issue is how multiple capabilities can be coordinated inside a single embodied intelligence system. The robot must see, understand, plan, act, and correct itself. No single trick can cover all of those requirements.

This is why the term embodied intelligence has become broader and more demanding. It is not simply about putting a large model inside a robot. It is about allowing a physical agent to function in an unstructured world. That world includes objects with mass, friction, fragility, and unexpected motion. It includes humans with unpredictable behavior. It includes tasks that unfold over long periods of time. An embodied intelligence system must survive contact with that reality.

2. The Debate Cools, but the Data Question Heats Up

Su Hang, an associate researcher in the Department of Computer Science and Technology at Tsinghua University, offered a widely cited judgment during a roundtable at the World Robot Conference. In his view, VLA and world models are not an either-or choice. For data that is high in quality and clearly labeled with actions and tasks, VLA can directly establish a mapping among perception, language, and action. For larger-scale and more diverse video data, a world model can help the system learn the relationship between environmental change and behavioral outcomes.

Su Hang predicted that the more likely future is a fused architecture. Such an architecture would combine the action-generation ability of VLA with the state-prediction and feedback-correction ability of a world model at different time scales. This is a significant statement for embodied intelligence because it moves the discussion away from ideological camps. It suggests that different model families may serve different roles in the same system. The question is not which one wins, but how they cooperate.

That shift is already visible in company practices. VLA has not been eliminated. Instead, it is being redefined. In many cases, VLA is no longer treated as a standalone technical route. It is becoming part of a broader system architecture. The robot still needs a way to connect what it sees and what it is told with what it does. That mapping remains essential. But mapping alone is not enough for robust embodied intelligence.

The data question follows naturally. If VLA depends on high-quality action and task labels, where do those labels come from? If world models depend on large-scale video and physical interaction data, how are those data collected, filtered, and aligned? The industry is discovering that data is not a secondary concern. It is the foundation on which every architecture claim must stand. Embodied intelligence cannot be built on model architecture alone.

This is why the 2026 conference felt cooler but more serious. The loudest arguments have faded. The harder work has begun. Companies are asking how to collect data in real environments, how to build training platforms, how to evaluate closed-loop performance, and how to ensure that a robot’s actions generate useful feedback. Those questions are less glamorous than a new model name, but they are more decisive for embodied intelligence.

3. VLA and World Models Move from Competition to Combination

Several companies at the World Robot Conference presented systems that combine VLA with world-model ideas. Keenon Robotics introduced its new-generation model KOM3.0, described as the world’s first service-industry VLA architecture that integrates a latent-space world model. The stated core breakthrough is that the robot rehearses physical outcomes internally before executing an action. It predicts the motion trajectory of an object after force is applied and the causal chain of multi-step operations. The goal is to avoid failure at the source rather than relying on trial-and-error execution.

The logic behind this approach is direct. VLA solves the problem of seeing, understanding, and working. A world model solves the problem of predicting consequences and reducing trial and error. When the two are layered, the robot can think before it acts. For embodied intelligence, that internal rehearsal is important because physical mistakes can be costly. A robot that can simulate the likely result of an action before taking it has a better chance of operating safely and efficiently.

Gao Jiyang, CEO of Galaxea, offered a more radical judgment. In his view, VLA and world models share the same core. Both are based on Transformer architectures and self-supervised pretraining. The difference lies mainly in the self-supervision method. He described the two as co-origin and co-growing, and predicted that they will eventually converge. This perspective challenges the idea that VLA and world models are separate species. Instead, it treats them as variations within a broader family of embodied intelligence architectures.

Yuandian Robotics, under Lexiang Technology, emphasized that a single model can hardly solve all problems of embodied intelligence in the real physical world. VLA addresses the mapping from vision and language to action. But in unstructured environments such as homes, additional capabilities are required. Data platforms, training platforms, and body systems must work together. The result is a continuously evolving capability loop. In this view, embodied intelligence is not a model file. It is a system that improves through interaction.

These examples show that the VLA versus world model framing is becoming obsolete. The more productive question is how to combine action generation with consequence prediction. A robot needs to act, but it also needs to anticipate. It needs to follow instructions, but it also needs to understand physical constraints. VLA and world models are increasingly seen as complementary components of embodied intelligence rather than competing final answers.

4. World Models Multiply, but Definitions Remain Fragmented

Everyone at the conference seemed to be talking about world models. Yet the models being built were not necessarily the same. The lack of a shared definition is one of the most important tensions in embodied intelligence today. A world model can mean a generative system that predicts future frames. It can mean a causal engine that predicts the consequences of actions. It can mean an event-level understanding system. It can also mean a latent-space simulator that supports planning without reconstructing every pixel.

Huang Yuanhao, founder of Orbbec, said during a main forum at the World Robot Conference that large language models can only address poetry and distant horizons, while world models are what allow robots to truly work. In his framing, GPT made human-machine conversation a reality. But the world model must address a much larger problem: the labor of eight billion people. The ambition is to liberate humans from repetitive labor. That ambition places world models at the center of embodied intelligence.

Hu Luhui, founder of Zhizheng AI, pointed to an even more fundamental issue. In his view, the world model must solve one of the core propositions of artificial intelligence: understanding the physical world. If multimodal large models remain at the level of perception, then world models must enter the deeper water of causal reasoning and decision-making. He breaks understanding the physical world into three levels. The first is cognition of the physical properties of objects. The second is cognition of the dynamic laws of environmental change. The third is task-oriented behavior planning.

For Hu Luhui, the world model is an indispensable underlying infrastructure for humanoid robots. A robot observes its surroundings and outputs its next action. To do so, it must rely on physical rules and dynamic patterns. It must make judgments through causal reasoning. Zhizheng AI has chosen the JEPA direction, or Joint Embedding Predictive Architecture, to learn physical laws in abstract representations. The company’s self-developed Chengling physical world model can learn latent representations of physical-world behavior from raw inputs such as vision and proprioception. This allows the robot to complete a mental simulation before executing a task.

Wujie Dynamics released MWA, described as the world’s first latent-space world model with a long-horizon bidirectional physical causal chain. It completes unified representation and a reasoning loop for multimodal information in latent space. Unitree focused on a video-generation path. Its WALL-WM model aligns vision, language, and action through events, allowing the robot to simultaneously understand the process, state changes, and results. These are different technical expressions of the same broad ambition: to give embodied intelligence a predictive inner world.

Jia Kui, founder of Crossdim Intelligent, described the world model as the crown jewel of generative AI and the endgame of AI unification. In his view, it can not only drive robots to work but also generate content whose three-dimensional physical geometry is absolutely correct. That is a strong claim. It connects world models not only to robotics but also to a broader generative AI future. For embodied intelligence, the promise is a system that understands physical structure well enough to generate and manipulate it.

Yet the challenges are real. Wang Xiaogang, chairman of Daxiao Robotics, said during the World Robot Conference that world models have already passed through three stages of evolution: data augmentation, simulation, and edge deployment. But they are still climbing. The leap from static scenes to dynamic real environments is far more difficult than expected. A world model that works in a controlled demonstration may fail in a noisy, changing, unpredictable physical space.

The greater challenge is that the industry has not unified its underlying definition of a world model. As Wang Xiaogang noted, what different teams are doing is still only roughly aligned at the bottom layer. Even evaluation standards are not fully unified. This fragmentation makes comparison difficult. It also makes it hard to know whether progress in one world model translates into progress for embodied intelligence as a whole.

5. Unified Representation Emerges as a New Route for Embodied Intelligence

Beyond the debate over VLA and world models, another route is emerging: the unified model. The Beijing Humanoid Robot Innovation Center introduced Pelican-Unify at the World Robot Conference, described as the country’s first unified-representation embodied world model. Its core breakthrough is that it no longer separates visual understanding, action control, and world prediction into independent modules. Instead, the three share one representation system and evolve together.

According to the center, whether the input is text, image, video, or action encoding, it is aligned into a shared representation space. Reasoning and planning are completed within the same framework. Actions are then generated end-to-end. The result is that understanding, reasoning, planning, and action are no longer serial stitches. They become an integrated closed loop. For embodied intelligence, this is a significant architectural claim because it attacks the problem of fragmentation directly.

Xiong Youjun, CEO of the Beijing Humanoid Robot Innovation Center, explained the distinction in an interview. Traditional VLM models can see and think but cannot act. VLA models can act but cannot rehearse. World models can rehearse but are not good at long-horizon reasoning. Pelican-Unify unifies vision, text, and action under the same paradigm. In May 2026, it topped WorldArena’s global authoritative evaluation ranking. The center presents this as evidence that embodied intelligence is moving from functional stitching to collaborative evolution.

This unified approach does not reject VLA or world models. It absorbs their concerns into a broader representation framework. The problem with separate modules is that errors can accumulate at boundaries. A perception module may produce one interpretation, a planning module may assume another, and a control module may execute something else. A unified representation aims to reduce those mismatches. It tries to make the robot’s internal world coherent.

For embodied intelligence, coherence is not an abstract virtue. A robot that sees a cup, plans to grasp it, and then applies force must maintain a consistent understanding of the cup’s position, shape, and fragility. If those representations are split across unrelated spaces, the risk of failure rises. A unified model attempts to keep perception, prediction, and action in the same conceptual currency. That is why it has attracted attention as a possible path beyond the VLA versus world model argument.

6. Brain-Inspired and Self-Evolving Architectures Challenge the Mainstream

Not every company is following the VLA-plus-world-model or unified-model path. Some are pursuing architectures inspired by biology. Zhipingfang introduced and open-sourced an original brain-inspired embodied model called NeuroVLA at the World Robot Conference. It brings the collaborative mechanism of the human brain’s cortex, cerebellum, and spinal cord into robot control systems. The cortex handles semantic understanding and task planning. The cerebellum handles high-frequency motor coordination and dynamic correction. The spinal cord handles millisecond-level action execution and safety reflexes.

Zhang Peng, co-founder of Zhipingfang, explained the logic behind this choice. In his view, VLA and world models are not a matter of choosing one path. They are part of a larger question: what kind of model should be used to enhance systemic capability? He noted that the company began working on end-to-end systems in 2023, introduced a fast-slow dual system in 2024, and launched NeuroVLA in 2026. The core goal is to make the system think and move more like a human. For embodied intelligence, this is an attempt to combine high-level reasoning with low-level reflex and control.

Shenyi Robotics proposed a different concept: self-evolution. Huang Kangyao, co-founder of Shenyi Robotics, said during a roundtable at the World Robot Conference that embodied scenarios exist in real physical space. Unlike large language models, where text can be traversed, real space requires a robot to output and generate continuous action control. This means data cannot be fully collected in advance. The company has proposed two dimensions internally: L4E, or Learning for Embodiment, meaning learn first and then work, and E4L, or Embodiment for Learning, meaning the robot continues learning while working in real tasks.

On the model route, Shenyi Robotics does not follow VLA or world models. It uses an original architecture internally called a hybrid physical expert model. The aim is to integrate physical-law understanding, motion prediction, obstacle avoidance, 3D estimation, and scene understanding into a single self-evolution paradigm. This is another expression of the same underlying search: how can embodied intelligence continue to improve when the world cannot be fully sampled, labeled, or simulated in advance?

These alternative routes matter because they expose the limits of any single architecture. A brain-inspired model may offer better real-time control. A self-evolving system may offer better adaptation. A unified representation may offer better coherence. None of these approaches can be dismissed simply because it does not use the most popular vocabulary. The field of embodied intelligence is still too young for final closure.

7. A Comparison of Embodied Intelligence Routes

The following table summarizes the major routes discussed at the 2026 World Robot Conference. It is not a ranking. It is a map of how different teams are approaching the robot brain. Each route reflects a different emphasis within embodied intelligence, and each faces open questions.

Route Representative Example Core Emphasis Stated Strength Stated Challenge or Open Question
VLA Keenon Robotics KOM3.0, described as a service-industry VLA architecture integrating a latent-space world model Mapping vision and language to action Directly connects perception, language, and action for labeled tasks Needs high-quality action and task labels; action mapping alone may not predict consequences
World Model Zhizheng AI Chengling, Wujie Dynamics MWA, Unitree WALL-WM Predicting physical outcomes, causality, and environmental change Supports internal rehearsal, state prediction, and feedback correction Definitions remain fragmented; evaluation standards are not unified; static-to-dynamic transfer is difficult
Fused VLA and World Model Keenon Robotics KOM3.0, Galaxea perspective, Yuandian Robotics system view Combining action generation with consequence prediction Allows the robot to think before acting and reduce trial and error Requires coordination across time scales and robust data pipelines
Unified Model Beijing Humanoid Robot Innovation Center Pelican-Unify Shared representation for vision, text, and action Aims for an integrated closed loop of understanding, reasoning, planning, and action Must prove coherence and scalability across diverse embodied intelligence tasks
Brain-Inspired Architecture Zhipingfang NeuroVLA Cortex, cerebellum, and spinal cord collaboration Combines semantic planning with high-frequency motor coordination and safety reflexes Complex to integrate; must demonstrate broad real-world reliability
Self-Evolving and Hybrid Expert Model Shenyi Robotics L4E and E4L, hybrid physical expert model Continuous learning in real physical work Targets adaptation where data cannot be fully collected in advance Requires long-term closed-loop deployment and stable self-evaluation

The table shows why the phrase embodied intelligence cannot be reduced to a single technical slogan. Each route addresses a real part of the problem. VLA addresses action mapping. World models address prediction and causality. Unified models address representation coherence. Brain-inspired models address control hierarchy. Self-evolving systems address adaptation over time. The future of embodied intelligence may not be a single route at all. It may be an integration of several routes into a practical system.

8. The Last Mile Runs Through Real-World Data and Closed Loops

A key observation from the 2026 World Robot Conference is that technical routes have not converged, but the direction is becoming clearer. The industry consensus is shifting away from the question of who owns the ultimate robot brain route. It is moving toward more practical questions. Whether a company chooses a fused VLA and world model route, a unified model route, a brain-inspired route, or a hybrid expert route, it must eventually return to the same starting point: data.

Embodied intelligence is moving from stages to homes and factories. That transition changes the standard. A robot that can move on a stage is not the same as a robot that can work in a factory. A robot that can run in a competition is not the same as a robot that can live in a family. The physical world is messy. It is full of variation, interruption, and unexpected events. The last mile from brain to body cannot be crossed by architecture alone.

Data must be collected in real scenes. It must be labeled with actions and tasks where possible. It must capture failures as well as successes. It must include long-horizon interactions, not just short demonstrations. It must support the training of VLA models, world models, unified models, and self-evolving systems. For embodied intelligence, data is not a static asset. It is a living part of the system. Every deployment can generate new information. Every correction can improve future behavior.

The closed loop is equally important. A robot must not only collect data but also act on it. It must perceive the result of its actions. It must compare predicted outcomes with actual outcomes. It must correct its internal model. This is where world models, VLA systems, and unified architectures meet the real test. A model that cannot close the loop in a real environment remains a laboratory artifact. A model that can close the loop becomes part of embodied intelligence.

The 2026 World Robot Conference did not announce a single winner. It did something more useful. It clarified the terms of the contest. VLA is not dead. World models are not a magic answer. Unified models are not automatically superior. Brain-inspired and self-evolving systems are not fringe curiosities. All of them are attempts to solve different pieces of the same puzzle. The puzzle is how to build embodied intelligence that can understand, predict, act, and improve in the physical world.

In the end, the strongest signal from the conference is practical. Technical route choice matters, but it is not the only thing that matters. The more decisive question is who can run data through real scenarios and complete a stable closed loop. That is the path from brain to body. That is the path from demonstration to deployment. That is the path that embodied intelligence must travel if it is to move from the stage to the home and from the laboratory to the factory floor.

As the industry continues to debate VLA, world models, unified architectures, and brain-inspired systems, the center of gravity is shifting. The next breakthroughs in embodied intelligence may come less from a new model name and more from the unglamorous work of data collection, alignment, evaluation, and closed-loop deployment. The robot brain will not be built by argument alone. It will be built by running the loop in the real world, again and again, until the machine can think before it acts and learn while it works.

Scroll to Top