I still remember the first time I watched a humanoid robot cross a laboratory floor. The robot moved slowly, arms slightly bent, feet lifting just enough to avoid tripping on a cable. It looked alive in a way that reminded me of a child learning to walk. Everyone in the room smiled. But when a visitor asked me whether this humanoid robot was “intelligent,” I could not give a clear answer. We had tested its motors, its batteries, its cameras, and its embedded computer, but we had no unified way to describe the intelligence of the whole machine. That question stayed with me. It became even more urgent as humanoid robot products began to appear from many companies, each claiming to be better, smarter, and more useful than the others. Without a common yardstick, the industry was speaking in dozens of different dialects.

Today, I believe we finally have the beginning of an answer. A new group standard for humanoid robot intelligence grading has been developed by a collaboration of innovation centers, major manufacturers, and research institutes, and it has been released by the China Electronics Society. The standard is numbered T/CIE 298-2025. It is the world’s first standard specifically designed for grading the intelligence of a humanoid robot. I was fortunate to be close enough to this process to watch the discussions, the debates, the technical drafting, and the final synthesis. This article is my personal reflection on why this humanoid robot intelligence grading standard matters, how it is structured, and what it may mean for the future of embodied artificial intelligence.
The Problem: A Humanoid Robot Has No Report Card
Let me start with the problem. In the past decade, I have seen humanoid robot demos that range from walking in a straight line to running, jumping, dancing, opening doors, picking up boxes, and even having short conversations with visitors. Yet almost all of those demos were evaluated in an informal way. Journalists wrote that a humanoid robot was “amazing,” engineers said that a robot was “stable,” and executives claimed that a robot was “world-leading.” These words are imprecise. When a humanoid robot falls during a performance, some people call it a failure while others call it a necessary experiment. When a humanoid robot successfully completes a task in a laboratory, some people call it general intelligence while others point out that the environment was carefully arranged. The industry needs a shared language to describe what a humanoid robot can actually do.
The automotive industry solved a similar problem with its levels of driving automation. People can discuss “Level 2” or “Level 3” driving assistance because there is an internationally recognized scale. In the same way, the new standard introduces a ladder for humanoid robot intelligence. Instead of vague adjectives, we can say that a humanoid robot operates at Level 3 in a specific scenario. Instead of arguing about demos, we can compare test results against a common set of indicators. This is the same kind of intellectual shift that the autonomous vehicle industry experienced, and I think the humanoid robot field is ready for that shift now.
The Birth of a Global First
I was not the only person who felt this urgency. The drafting of the humanoid robot intelligence grading standard began as a collective effort. The leading innovation center and its partner organizations in different regions worked together with manufacturers, public institutions, and academic teams. These organizations contributed data from real humanoid robot products, including bipedal robots, wheeled humanoid machines, and upper-body humanoid systems. The goal was to create a framework that could be applied to any humanoid robot, regardless of its mechanical design or its target market.
What emerged is called a “four-dimensional five-level” evaluation framework. The four dimensions are perception and cognition, decision and learning, execution and performance, and collaboration and interaction. Each dimension has a symbol: \(P\) for perception and cognition, \(D\) for decision and learning, \(E\) for execution and performance, and \(C\) for collaboration and interaction. The five levels are \(L1\) through \(L5\). According to the standard, a humanoid robot at \(L1\) demonstrates the lowest level of intelligent capability, while a humanoid robot at \(L5\) demonstrates the highest level. From \(L1\) to \(L5\), the humanoid robot’s ability to perceive its environment, make decisions, execute actions, and work with people increases step by step.
This standard is not just a simple checklist. It contains 22 first-level indicators, more than 100 technical clauses, a common safety baseline, and mappings to typical application scenarios. For product designers, it provides a guide for what to measure. For buyers, it provides a way to compare competing humanoid robot products. For researchers, it provides a common vocabulary to describe experimental results. For regulators, it provides a foundation for certification and market supervision. I believe this is exactly what the humanoid robot industry needs at this moment: not another promise, but a transparent system of evaluation.
The Four Dimensions of Humanoid Robot Intelligence
Let me explain the four dimensions in more detail. When we discuss a humanoid robot, we are not discussing a pure algorithm that lives on a server. A humanoid robot has a body that moves through the physical world. It has sensors that must interpret noisy, incomplete, and constantly changing information. It has a brain, or a control stack, that must decide what to do next. It has actuators that must express those decisions in muscle-like forces. And it has to do all of this while sharing space with people. The four dimensions capture these requirements.
| Dimension | Symbol | What It Measures | Examples of Evaluation Subjects |
|---|---|---|---|
| Perception and Cognition | \(P\) | The ability to sense the world, understand scenes, recognize objects, and build an internal model of the environment. | Object detection, semantic segmentation, human pose estimation, depth estimation, state estimation of the robot’s own body. |
| Decision and Learning | \(D\) | The ability to plan actions, reason about tasks, learn from experience, and adapt to new situations. | Task decomposition, motion planning, skill learning, reinforcement learning, failure recovery planning. |
| Execution and Performance | \(E\) | The ability to move, balance, manipulate objects, sustain operation, and complete physical tasks with precision and stability. | Walking speed, bipedal stability, stair climbing, obstacle crossing, grasping success rate, payload capacity, endurance. |
| Collaboration and Interaction | \(C\) | The ability to work with humans, communicate through language and gestures, understand intent, and follow social rules. | Natural language understanding, speech generation, gesture recognition, joint task execution, safe human-robot handovers. |
For a humanoid robot, these four dimensions cannot be separated too sharply. Perception informs decisions; decisions are useless without execution; and execution gains meaning only when the robot collaborates with people. In my own experience with humanoid robot testing, a robot can have excellent perception and still fail because its legs are unstable. Another humanoid robot may have brilliant decision algorithms but fail to finish a task because its hands cannot grasp tools effectively. Therefore, the standard treats the four dimensions as a combined evaluation system rather than four separate tests.
A Simple Model of Overall Intelligence
How does one combine the four dimensions into a single grade? One possible approach is a weighted sum. In very general terms, we can define an overall intelligence score for a humanoid robot as:
\[
I_{\text{robot}} = w_P P + w_D D + w_E E + w_C C
\]
Here, \(P\), \(D\), \(E\), and \(C\) are the normalized scores for the four dimensions, and \(w_P\), \(w_D\), \(w_E\), and \(w_C\) are weights that reflect the importance of each dimension in a given scenario. The weights should satisfy:
\[
w_P + w_D + w_E + w_C = 1, \qquad w_P, w_D, w_E, w_C \ge 0
\]
In practice, the standard does not simply average the four dimensions. A humanoid robot that scores very high in perception but dangerously low in execution should not receive a high grade. In my conversations with other engineers, we often discussed the idea of a “bottleneck.” A robot is only as intelligent as its weakest component in the physical world. If it can think faster than it can move, the body becomes the constraint. If it can move faster than it can think, the control system becomes the constraint.
For this reason, I find it useful to think about a humanoid robot’s effective capability as a combination of the four dimensions with a bottleneck constraint. A simple illustration is:
\[
I_{\text{effective}} = \min\left( \frac{P + D}{2}, E, C \right)
\]
This is not the official formula of the standard; it is a mental model that I use to explain why a humanoid robot with balanced capabilities is often more impressive than one with a single very strong ability. In the end, the official standard uses a more rigorous mapping to determine levels, but the principle remains: the level of a humanoid robot is determined by the whole system, not by one marketing story.
The Meaning of L1 to L5
The five levels in the humanoid robot intelligence grading standard are designed to be cumulative. Each higher level represents a meaningful increase in autonomous capability. The standard does not define these levels as simple percentages, because a humanoid robot’s capabilities are too diverse. Instead, it describes the level of autonomy, the degree of human supervision, the complexity of the environment, and the difficulty of the tasks that a humanoid robot can handle.
| Level | General Concept | Level of Human Supervision | Typical Task Environment |
|---|---|---|---|
| \(L1\) | Basic assisted operation: the humanoid robot performs motions under continuous human guidance or teleoperation. | Continuous human control | Fixed structures, controlled laboratories, remote hazardous sites |
| \(L2\) | Task-specific autonomy: the humanoid robot can execute specific tasks automatically while a human supervises. | Human approval for key decisions | Structured factory stations, simple pick-and-place workflows |
| \(L3\) | Conditional autonomy: the humanoid robot handles routine operations in constrained scenes and asks for help when uncertain. | Occasional human takeover | Warehouse logistics, industrial inspection, indoor service settings |
| \(L4\) | High autonomy: the humanoid robot plans, acts, and adapts within a defined scenario with broad robustness. | Remote monitoring and exception handling | Hospital corridors, retail stores, factory complexes, outdoor pedestrian areas |
| \(L5\) | General autonomy: the humanoid robot performs novel tasks in open environments with human-like reasoning and social awareness. | Minimal supervision | Households, unknown buildings, disaster response, long-horizon mission spaces |
From \(L1\) to \(L5\), we can see the path of evolution. A humanoid robot at \(L1\) is mostly a remote-controlled machine. A humanoid robot at \(L3\) can work in limited places without being constantly guided. A humanoid robot at \(L5\) would approach what many people call artificial general intelligence embedded in a physical body. The standard does not promise that all humanoid robot products will reach \(L5\) soon. Instead, it gives the industry an honest map of the road.
In the standard, the final grade may be represented by a threshold function. For each dimension \(d\), there is a threshold \(\theta_{d,l}\) for each level \(l\). The humanoid robot can claim level \(l\) only if its score in every relevant dimension satisfies the corresponding threshold. A conceptual way to write this is:
\[
L_{\text{humanoid}} = \min_{d \in \{P, D, E, C\}} \max\left\{ l \in \{1,2,3,4,5\} : S_d \ge \theta_{d,l} \right\}
\]
This formula expresses a strict view of intelligence grading: the humanoid robot is no stronger than its weakest dimension. I think that is the right approach. A robot that talks beautifully but cannot carry a box is not ready for the physical world. A robot that carries boxes brilliantly but cannot understand a simple instruction is not ready for collaboration.
Indicators and Technical Clauses
The humanoid robot intelligence grading standard includes 22 first-level indicators and more than 100 technical clauses. These clauses translate abstract concepts into measurable tests. Instead of saying “the robot should be aware of its environment,” the standard contains clauses about object detection precision, scene classification accuracy, and response time. Instead of saying “the robot should walk well,” the standard asks about walking speed, step height, incline tolerance, recovery from pushes, and the robot’s ability to remain upright after disturbances.
| Dimension | Example First-Level Indicators | Example Technical Clauses |
|---|---|---|
| \(P\) | Object recognition, scene understanding, self-state estimation, human perception | Detect a chair within 0.5 meters at 95% confidence; recognize a falling person; estimate the robot’s pelvis position at 50 Hz. |
| \(D\) | Task planning, skill learning, motion planning, adaptive decision-making | Generate a multi-step task plan from a natural language instruction; recover from a failed grasp within two seconds; learn a new manipulation skill within ten demonstrations. |
| \(E\) | Locomotion stability, dynamic balance, manipulation dexterity, endurance | Walk on uneven terrain without stepping outside a defined corridor; climb stairs of 15 cm height; lift a package of 5 kg and place it on a shelf. |
| \(C\) | Natural language interaction, gesture communication, social compliance, team collaboration | Answer a human request using spoken language; follow a pointing gesture; stop immediately when a person enters the safety zone; coordinate with another robot to carry a long object. |
Having worked with humanoid robot testing, I know how difficult it is to define these clauses in a way that is fair across different designs. A small humanoid robot cannot lift the same load as a large industrial humanoid robot. A performance-focused humanoid robot cannot be judged solely by its conversational ability. Therefore, the standard includes typical application scenarios. The same humanoid robot can be graded differently for different application scenarios. For example, a humanoid robot used in education might emphasize \(P\) and \(C\), while a humanoid robot used in logistics might emphasize \(E\) and \(D\). The overall grade is therefore not a single permanent label but a profile that should be interpreted together with the intended application.
Safety as an Unbreakable Gate
No humanoid robot should receive a high intelligence grade if it is unsafe. The standard includes a universal safety bottom line that applies to every level. This is one of the most important features, in my opinion. A humanoid robot is a physical system that can move quickly, carry heavy objects, and work alongside people. If its intelligence is not accompanied by reliable safety mechanisms, it should not be deployed outside a test environment. The safety baseline includes requirements for emergency stopping, collision avoidance, force limiting, risk assessment, fail-safe behavior, and transparent reporting of limitations.
In a formal sense, the safety gate can be represented as follows. Let \(S_{\text{safety}}\) be the robot’s aggregate safety score, and let \(S_{\min}\) be the minimum threshold defined by the standard. The final grade \(G_{\text{final}}\) is:
\[
G_{\text{final}} =
\begin{cases}
L_{\text{humanoid}}, & \text{if } S_{\text{safety}} \ge S_{\min},\\
\text{not graded}, & \text{otherwise}.
\end{cases}
\]
This is not a minor detail. During the drafting discussions, I heard many vivid stories about humanoid robot accidents: a swinging arm, a sudden fall, a false detection, a robot that knocked over a display stand. Those incidents were not necessarily signs that the robots were badly designed; they were signs that the industry needed a common language for acceptable risk. The standard gives manufacturers a clear target: if you want to claim a level for your humanoid robot, you must first pass the safety gate.
How a Score Could Be Computed
Let me explain one way that the indicators could be normalized and combined. For each technical clause \(i\) in dimension \(d\), the test produces a raw measurement \(x_{d,i}\). The standard defines a minimum score \(x_{d,i}^{\min}\) and a maximum score \(x_{d,i}^{\max}\). A normalized score for that clause can be computed as:
\[
\bar{r}_{d,i} = \frac{x_{d,i} – x_{d,i}^{\min}}{x_{d,i}^{\max} – x_{d,i}^{\min}}
\]
For a binary clause, such as whether the humanoid robot can stop in response to an emergency signal, the normalized score is either 0 or 1. For a continuous clause, such as walking speed or grasping success rate, the normalized score can take any value between 0 and 1. The dimension score \(S_d\) is the weighted sum of the clause scores within that dimension:
\[
S_d = \sum_{i=1}^{N_d} \lambda_{d,i} \bar{r}_{d,i}
\]
where \(N_d\) is the number of technical clauses in dimension \(d\), and \(\lambda_{d,i}\) is the weight of clause \(i\). The weights within each dimension satisfy:
\[
\sum_{i=1}^{N_d} \lambda_{d,i} = 1, \qquad \lambda_{d,i} \ge 0
\]
I want to emphasize that this is a simplified representation of the standard’s evaluation logic. The actual standard contains more nuance, especially when a single clause is considered critical. But the general idea is clear: a humanoid robot’s intelligence grade should be based on reproducible measurements, not on subjective impressions.
An Illustrative Comparison of Two Humanoid Robots
To see how this scoring framework might help, imagine two humanoid robot prototypes shown in the same exhibition. The first humanoid robot walks beautifully. It does not fall during a ten-minute demonstration. It climbs a gentle slope and steps over a small obstacle. However, when a visitor asks it a question, it shows no response. The second humanoid robot talks confidently with visitors and answers questions about the weather, but it remains mounted on a stand and cannot walk at all. Which humanoid robot is more intelligent? The answer depends on the weights chosen for the task and scenario. In a mobility-oriented mission, the first robot is more useful. In a reception-oriented scenario, the second robot is more useful. The standard forces us to make these weights and measurements explicit.
| Humanoid Robot Prototype | \(P\) Score | \(D\) Score | \(E\) Score | \(C\) Score | Weighted Overall Score |
|---|---|---|---|---|---|
| Prototype A: walking and balancing focus | 2.8 | 2.5 | 4.3 | 1.4 | 2.78 |
| Prototype B: conversation and interaction focus | 3.8 | 3.5 | 1.2 | 4.4 | 3.33 |
In this illustrative table, Prototype B might receive a higher overall score if the weights favor communication and reasoning. Prototype A might receive a higher grade in a logistics scenario that emphasizes execution performance. The standard does not hide these trade-offs. It makes them visible. This is precisely why I believe the humanoid robot grading standard will change how companies design and market their products.
Policy and Standardization Roadmap
The emergence of this standard was not accidental. It is part of a broader policy push to make the humanoid robot industry more mature and orderly. I remember when the central guidance on humanoid robot innovation was published in 2023. That document explicitly asked for a standardization roadmap, a review of the industrial chain, and a systematic approach to standards. At the time, I thought the idea was ambitious. Now I see that it has led to concrete results.
| Year | Milestone | Impact on Humanoid Robot Industry |
|---|---|---|
| 2023 | A national innovation deployment plan for humanoid robots was issued. | It encouraged the creation of standards and the definition of industrial categories. |
| 2024 | A humanoid robot standardization technical committee was planned and publicly announced. | It created an institutional home for future humanoid robot standards. |
| 2025 | The draft of the intelligence grading standard was released for public comment in February. | Manufacturers and researchers could review and submit feedback before the final text was issued. |
| 2025 | The standard T/CIE 298-2025 was officially released. | It gave the humanoid robot industry a common language for intelligence levels. |
I was particularly interested in the public comment phase. When the draft appeared, our lab set aside a week to read every clause and ask hard questions. Did a clause make sense for a wheeled humanoid robot? Could a small humanoid robot ever satisfy the same execution thresholds as a large one? How should the test environment be standardized? Many of these questions were debated in the comments. In the final version, I believe the standard succeeded in balancing comprehensiveness and flexibility.
From Mechanical Fantasy to a Golden Track
Let me zoom out and look at the broader history. The first humanoid robot was created in 1973. It was a remarkable mechanical feat for its time, but its intelligence was minimal. Since then, the development of humanoid robots has gone through several phases. From the 1970s to the 2000s, the focus was on mechanical anthropomorphism: legs, arms, head, and joints. From roughly 2010 to 2020, the humanoid robot field entered a period of dynamic control, in which robots learned to run, jump, climb, and maintain balance under perturbations. From 2020 to 2025, artificial intelligence began to change the picture. A humanoid robot could not only move like a person; it could also perceive, plan, and learn in more sophisticated ways.
| Period | Phase Name | What a Humanoid Robot Could Do |
|---|---|---|
| 1970s-2000s | Mechanical anthropomorphism | Walk slowly, move arms, replay preprogrammed motion sequences, demonstrate basic interaction. |
| 2000s-2010s | Whole-body coordination | Keep balance, climb stairs, avoid obstacles, perform human-like gestures, respond to simple commands. |
| 2010s-2020 | Dynamic control | Run, jump, recover from pushes, perform acrobatic motions, use force control for manipulation. |
| 2020-2025 | Artificial intelligence empowerment | Understand scenes, learn new tasks, adapt to changing environments, interact through language, work in real applications. |
I have seen this shift with my own eyes. In the early 2020s, a humanoid robot might fall several times during a short walk. By the middle of the decade, I was watching humanoid robots play football, stand up after being tackled, climb long staircases, run in snow, and remain upright even when people pushed them. In one memorable event, a humanoid robot ran a half-marathon alongside human athletes. Some robots fell, but what impressed me was not perfection; it was the speed of improvement. Inside the industry, people often say that almost every week brings a small breakthrough. That may sound like an exaggeration, but after watching the last few years, I believe it is close to the truth.
The development of humanoid robot intelligence can be approximated by a growth curve. If we define \(C(t)\) as the average capability of a humanoid robot at time \(t\), then the rate of improvement has accelerated dramatically in recent years:
\[
\frac{dC(t)}{dt} = \beta C(t)
\]
where \(\beta\) is the growth coefficient. In the early decades, \(\beta\) was small because progress depended on mechanical design and manual tuning of control systems. After the introduction of deep learning, simulation, and large-scale models, \(\beta\) increased sharply. The humanoid robot field moved from a slow accumulation of hardware improvements to an exponential improvement of intelligence and behavior.
The New Standard Arrives at the Right Time
Why is the new standard so important now? Because the humanoid robot field is entering a stage where products are no longer just academic curiosities. They are beginning to work in factories, warehouses, pharmacies, and public spaces. In an automobile factory, I have seen plans to use dozens of humanoid robots for material handling and assembly assistance. In a pharmacy, I have seen a humanoid robot autonomously pick medicine from shelves, place it in a bag, and hand the bag to a courier. In a research laboratory, I have seen humanoid robots learn to open doors, plug cables, and move unknown objects. Each of these applications requires a different mix of perception, decision, execution, and collaboration skills. Without a grading system, it is almost impossible for a customer to know whether a particular humanoid robot is suitable for a particular job.
The standard also addresses the problem of exaggerated claims. In the past, companies could say that their humanoid robot was “the most advanced in the world” without offering any reproducible evidence. With the new standard, they can be asked: “What level is your humanoid robot according to T/CIE 298-2025? In which dimensions did it pass? What safety baseline did it meet?” This kind of verification is essential for market trust. A humanoid robot is far too expensive and important to be purchased on the basis of a polished video alone.
Let me mention the concept of “display intelligence.” A humanoid robot may look brilliant during a carefully choreographed performance, but fail as soon as the environment changes. The new grading standard pushes the industry toward “general intelligence.” A humanoid robot should be able to understand the task, adapt to new situations, make sensible decisions, and safely collaborate with people. This is a much higher bar than performing a memorized sequence of motions. The standard does not magically create general intelligence, but it gives us a way to measure the distance between display intelligence and true capability.
Mapping Levels to Application Scenarios
The standard includes mappings from intelligence levels to typical application scenarios. These mappings help customers choose the right humanoid robot for a specific environment. The table below is my interpretation of how such mappings can be used.
| Application Scenario | Primary Humanoid Robot Capabilities | Typical Target Level |
|---|---|---|
| Special operations and hazardous environments | Execution, perception, remote control, autonomous recovery | \(L1\) to \(L3\), depending on telepresence and autonomy |
| Logistics and material handling | Execution, decision-making, navigation, manipulation | \(L3\) to \(L4\) |
| Industrial manufacturing | Execution, collaboration, precision, safety | \(L3\) to \(L4\) |
| Education and scientific research | Decision, perception, collaboration, modular programming | \(L2\) to \(L4\) |
| Commercial service | Collaboration, interaction, communication, mobility | \(L3\) to \(L4\) |
| Health care and elderly care | Perception, collaboration, interaction, safety | \(L4\) to \(L5\) |
| Residential and household service | All dimensions, especially open-world adaptation | \(L4\) to \(L5\) |
I find this mapping very useful because it shows that a humanoid robot does not need to be \(L5\) to create value. A humanoid robot that works well in a warehouse at \(L3\) may be deployed today. A humanoid robot that reliably assists nurses in a hospital at \(L4\) may be possible soon. A humanoid robot that can live in a family home and handle endless unpredictable tasks may wait until \(L5\) technology is mature. The standard allows the market to make these distinctions.
Cost, Localization, and Mass Production
The standard alone will not produce the future. I have spoken with robot executives and engineers who believe that the humanoid robot industry is still too early. They say that the cost of advanced sensors and actuators is too high, that the manufacturing volume is too low, and that many technical routes remain uncertain. I agree with those concerns. But I also believe that standards can help reduce cost. When a humanoid robot can be measured, compared, and specified, customers can make more confident purchasing decisions. Greater demand leads to larger production volumes. Larger production volumes lead to lower costs. Lower costs expand the market further.
The new standard encourages domestic production and localization in a subtle way. By defining a common baseline, it helps suppliers of motors, harmonic drives, sensors, artificial intelligence chips, and communication modules align their products with the needs of humanoid robot makers. In the coming years, I expect to see more core components designed specifically for humanoid robot applications. This will not happen overnight. It will require steady investment, engineering discipline, and long-term patience. But the standard gives the supply chain a clearer target.
I recall one conversation with a robot company founder who said that industrial scenarios are only a bridge. The real explosion of humanoid robot applications will occur in the home. The same person warned that the industry should prepare for a five-year or even ten-year marathon. Some technical routes have not yet been confirmed, and companies should not be distracted by short-term hype. What matters most is reducing cost, increasing production volume, and localizing the supply chain. I think the intelligence grading standard contributes to this long-term goal by creating a stable vocabulary. It helps everyone remember that the humanoid robot is not a magic product; it is an engineered system that can be improved step by step.
Government Support and the Road to Mass Deployment
Regional governments have also recognized the importance of humanoid robots and embodied intelligence. In 2025, an action plan was announced by the Beijing municipal government to nurture the next stage of humanoid robot and embodied intelligence development. The plan sets targets for 2027. I see these targets as a natural companion to the intelligence grading standard. The standard defines what level means; the action plan defines what the industry should achieve.
| Target Area | 2027 Goal |
|---|---|
| Key technologies | Break through at least 100 key technologies related to embodied intelligence and humanoid robots. |
| Flagship products | Release at least 10 internationally leading software and hardware products. |
| Industrial chain | Localize the upstream and downstream industrial chain of embodied intelligence. |
| Core enterprises | Cultivate at least 50 core enterprises in the upstream and downstream sectors. |
| Mass-production products | Form mass production of at least 50 humanoid robot products. |
| Application cases | Achieve at least 100 large-scale applications in science and education, industry and commerce, and personalized services. |
| Production volume | Reach a total humanoid robot production volume of more than 10,000 units. |
| Industrial scale | Build a humanoid robot and embodied intelligence industry cluster valued at 100 billion yuan. |
I am particularly encouraged by the commitment to create industrial parks for humanoid robot companies. One park is planned in the northern part of Beijing, while another already exists in the southern development area. These two parks can form a corridor that provides companies with research, development, manufacturing, testing, and application space. The new intelligence grading standard can serve as a testing and certification framework for these parks. When companies move into such an ecosystem, they can rely on the standard to evaluate their products, improve their designs, and communicate with investors and customers.
The Humanoid Robot and the Future of Work
What does the grading standard mean for the future of work? I think the answer lies in the word “collaboration.” A humanoid robot is not meant to simply replace a human worker; it is designed to work alongside people, using tools designed for human hands, and navigating spaces designed for human bodies. In a factory, a humanoid robot can take over repetitive, physically demanding, or dangerous tasks. In a hospital, it can carry supplies, assist with patient mobility, and communicate with care teams. In a logistics center, it can handle items that are too heavy for a human to lift repeatedly. In a home, it can help with cleaning, cooking, monitoring safety, and providing companionship. None of these applications is science fiction. They are already being tested.
The intelligence grading standard will help employers evaluate whether a humanoid robot is ready for a particular task. Instead of asking a vague question like “Can this robot work in my warehouse?” they can ask “Does this humanoid robot meet \(L3\) requirements for logistics and material handling? Does it have a sufficiently high execution score? Does it satisfy the safety baseline for working near humans?” This is a much more useful procurement process. It turns the humanoid robot market into a more rational market.
Of course, there are limits to what any test can predict. A humanoid robot may pass a certification test in a laboratory and then encounter an unexpected environment on the factory floor. That is why I believe the standard should be updated continuously. The humanoid robot industry is moving quickly, and the standard should move with it. It should not be treated as a final answer but as a living framework. In that sense, the first version of the standard is not the end of the conversation; it is the beginning.
My Concerns and Hopes
Let me be honest about my concerns. A grading standard can be abused. A manufacturer might optimize its humanoid robot to pass narrowly defined tests while ignoring the real-world diversity of tasks. A marketer might claim a level that applies only to a specific configuration, not to the product that customers actually receive. An engineer might spend too much time improving test scores and too little time improving the humanoid robot’s ability to handle unknown situations. I think the standard can help avoid these problems if the industry treats it with integrity. The standard should be used as a communication tool, not as a weapon for misleading advertising.
But I also have hope. In my years of working with humanoid robot systems, I have rarely seen an idea that made the industry feel more together. The intelligence grading standard gives us a shared map. It allows a researcher in one city to compare results with a researcher in another city. It allows a startup company to demonstrate its progress against an official scale. It allows a customer to specify needs in the language of the standard: “I need a humanoid robot with \(P4\) perception, \(D3\) decision-making, \(E4\) execution, and \(C3\) collaboration.” This kind of specification is not possible with adjectives alone.
The image of a humanoid robot working in a household, helping an elderly person, or assisting a nurse in a hospital is no longer a fantasy. The technology is advancing. The standard will help that technology leave the stage and enter reality. I have no doubt that the first version of this standard will need improvement. New tests will be invented. New application scenarios will appear. New artificial intelligence methods will make current thresholds seem too conservative or sometimes too generous. That is natural. What matters is that the humanoid robot field now has a foundation on which to build.
From Display Intelligence to General Intelligence
I want to return to the phrase “display intelligence.” A humanoid robot can be programmed to show off impressive skills for a few minutes. It can walk, wave, answer a prepared question, and shake hands with an audience member. But these behaviors do not necessarily mean that the humanoid robot understands the world. The new standard challenges the industry to move from display intelligence to something deeper: the ability to perceive, decide, execute, and collaborate in real time, under uncertainty, and across different environments.
This shift is similar to what happened in the autonomous vehicle industry. Early demonstrations of self-driving cars were exciting, but they were not enough to guarantee safety or scalability. The industry needed standardized levels of automation, clear definitions of driver supervision, and rigorous test protocols. The humanoid robot industry needs the same. The intelligence grading standard is the first major step in that direction. It will not solve every problem, but it will give engineers and customers a common point of reference.
I have come to believe that the humanoid robot is not just another product. It is a mirror for our own intelligence. By trying to build a machine that can walk, see, think, speak, and act in the world, we are forced to understand those abilities more deeply. The grading standard forces us to be precise about that understanding. It gives us a structure for asking: What does it mean for a humanoid robot to perceive? What does it mean for it to decide? What does it mean for it to learn? These are not only engineering questions; they are philosophical questions as well. The standard answers them with practical definitions and measurable criteria.
The Road Ahead for Humanoid Robot Intelligence
Where will the humanoid robot field be in 2027? If the action plan succeeds, there will be many more mass-produced humanoid robots in factories, schools, hospitals, and service centers. The intelligence grading standard will be used to certify those robots. I expect to see more detailed application-specific grading. A humanoid robot built for warehouse logistics might receive one level for manipulation and another level for navigation. A humanoid robot built for elderly care might receive a high level in collaboration but a lower level in heavy-load execution. These profiles will be more informative than a single letter grade.
I also expect to see more automated testing environments. Instead of humans manually measuring walking speed and grasping success, test facilities will use motion capture, sensor networks, and cloud data collection to evaluate humanoid robot behavior. The standard can support such automation because it defines what should be measured. A humanoid robot will enter a test course, perform tasks for hours, and leave with a detailed report of its scores in the four dimensions.
In the long run, I hope to see international collaboration on humanoid robot intelligence grading. The first standard was developed in China, but the need is global. A humanoid robot made in one country should be able to be compared with a humanoid robot made in another country, at least in terms of general intelligence levels. International cooperation will require careful adaptation, but the existence of a national standard is exactly the right starting point. We need a snowball, not a blueprint. The new standard is our snowball.
A Personal Conclusion
When I look at the image of a humanoid robot on the shelf of my laboratory, I no longer feel that I am looking at a purely mechanical object. I see a system whose intelligence can be measured, improved, and compared. The new standard has changed the way I think about my own work. Every time I watch a humanoid robot move, I now ask: What is its \(P\), \(D\), \(E\), and \(C\) profile? What level of autonomy does it truly have? Has it passed the safety gate? These questions sharpen my judgment.
The arrival of the humanoid robot intelligence grading standard means that the era of empty claims is coming to an end. From now on, a humanoid robot will be judged not by the beauty of its demo, but by the depth of its intelligence. This is exactly what the industry needs as it moves from mechanical fantasy to a golden track of practical applications. A humanoid robot will not become useful simply because it looks human; it will become useful because it can perceive, decide, execute, and collaborate. The standard gives us the scale to see how far we have come and how far we still have to go.
\[
I_{\text{humanoid}} = \sum_{d} w_d S_d, \qquad \text{with safety as the gate}
\]
In the end, I believe the humanoid robot industry is on the edge of a transformation. The next few years will determine whether humanoid robots become everyday tools or remain laboratory curiosities. The intelligence grading standard is not a magic solution. It cannot make a weak humanoid robot strong. But it can make progress visible. It can make evaluation fair. It can make claims honest. And it can give all of us who work on humanoid robot technology a clearer path toward the intelligent machines we have imagined for so long.
For me, the standard also settles an old personal question. When someone asks, “Is that humanoid robot intelligent?” I can now answer with more than an opinion. I can point to the dimensions, the levels, the indicators, the test clauses, and the safety baseline. I can say exactly what the robot can do and what it cannot do. That is the power of a good standard. It does not end the quest for intelligence; it begins a new and better chapter in that quest. The humanoid robot will not be judged by its shape, but by its mind. And now, at last, we have a ruler to measure that mind.
