In recent years, artificial intelligence research has undergone a deep transformation. The classical view of intelligence as abstract computation over symbolic representations has been challenged by the insight that cognition is fundamentally rooted in the body. In this essay, I reflect on how artificial intelligence can become embodied, using the humanoid robot as my central case. I argue that the humanoid robot is not merely a technological product but a philosophical provocation that forces us to reconsider the relation between body, mind, and world. My inquiry proceeds from a simple but radical question: what does it mean for an artificial system to possess a body in a way that is constitutive of intelligence, rather than merely instrumental to it? By examining the rise of embodied intelligence, the three dimensions that structure it, and the specific case of humanoid robots, I hope to show that the deepest obstacles to truly intelligent machines are no longer purely computational, but phenomenological and ethical.

1. The Paradigm Shift: From Disembodied Computation to Embodied Intelligence
The classical tradition in artificial intelligence, which dominated the field for decades, treats intelligence as a formal symbol-manipulation process. According to this view, thinking is equivalent to computing: the mind is a digital computer, and cognition consists of operations on internal representations that mirror the external world. In this framework, the body is reduced to a peripheral device that supplies sensory inputs and executes motor outputs. The essential locus of intelligence is a central processing unit, whether it is made of neurons or silicon. This set of assumptions, often called computationalism, representationalism, or cognitivism, has been the target of repeated critique since the late twentieth century. I have come to believe that these critiques are not merely philosophical quibbles, but point to a fundamental error at the heart of the classical project.
The first generation of AI systems achieved remarkable successes in formal domains such as chess, logic, and theorem proving. Yet these systems failed spectacularly when confronted with the messiness of everyday perception, flexible motor control, and social interaction. A robot that could calculate the trajectory of a falling object with perfect accuracy would still stumble if it had to reach out and catch the object in a cluttered, unpredictable environment. Why? Because such tasks require the continuous tuning of a body to the environment, a kind of intelligent responsiveness that is not captured by explicit rules and static representations. The failure of classical robotics to produce robust, flexible behavior in open environments led me, along with many others, to embrace a different starting point: intelligence must be understood as an emergent property of a living, moving, situated body.
This shift is often described as a move from the computational model toward an embodied, enactive, and ecological view. In a phenomenological framework, the body is not a shell that contains a mind; it is the very medium through which the world appears. The French phenomenological tradition has long taught that perception is not a passive reception of stimuli but an active, skill-laden engagement with the world. When I walk through a room, I do not first construct a mental model of the chairs, tables, and walls, then compute a plan of action. Instead, my body immediately grasps the affordances of the environment: this chair is felt as “graspable,” that corridor as “passable.” This pre-reflective, non-representational understanding is the foundation of all higher cognition. For artificial intelligence to be genuinely intelligent, it must somehow participate in this kind of embodied coupling, not merely simulate it in a symbolic model.
I therefore see the recent turn toward embodied intelligence as a return to a deeper truth, one that was already present in the phenomenological tradition long before artificial intelligence existed. The challenge is to translate this philosophical insight into a concrete research program. How can an artificial system—a robot, an agent, a machine—exhibit the kind of embodied intelligence that humans display effortlessly? The humanoid robot is a particularly fascinating case because it attempts to replicate the human body in its external form. But does a human-like form guarantee human-like embodied cognition? This is the question that I will return to throughout this article.
2. From Non-Embodied to Embodied: My Intellectual Journey
When I first studied artificial intelligence, the dominant paradigm was still the physical symbol system hypothesis. I was taught that the essential feature of an intelligent system is its ability to manipulate symbols according to rules. The physical realization of those symbols—whether in the brain, in a computer, or in a robot—was irrelevant. This doctrine, known as multiple realizability, seemed to make the body entirely accidental. If intelligence is defined at the level of algorithms, then a mind could exist equally in a human brain, a silicon chip, or, conceivably, in a sufficiently complex network of water pipes. I found this idea intellectually seductive but increasingly unsatisfying as I encountered the failures of rule-based systems in real-world contexts.
A turning point came when I read about pattern recognition. Classical attempts to program computers to recognize objects, faces, and speech were based on feature extraction, template matching, and statistical classification. These systems worked well in controlled environments but broke down in noisy, ambiguous, and novel contexts. Human pattern recognition, by contrast, is wonderfully robust. I can recognize a friend even in a crowd, from a strange angle, in poor lighting, when she has changed her hairstyle, and even when I am tired or distracted. This robustness is not simply a matter of better algorithms; it arises from a deep coupling between perception and action. I can move my head, squint, shift my gaze, walk around the object, and use my body to disambiguate what I see. In other words, my perceptual experience is not a static input-output computation, but a dynamic process of exploration and prediction that involves my whole body.
Another decisive influence was the work on reactive and behavior-based robotics. Some roboticists in the 1980s and 1990s argued that intelligent behavior could be generated by simple, layered control systems directly coupled to sensors and actuators, without any central representation or plan. Their robots could navigate around obstacles, explore rooms, and even exhibit something like goal-directed behavior, all without an internal model of the world. These experiments were not just engineering feats; they were existence proofs that intelligence does not require the kind of explicit representation that classical AI insisted upon. They demonstrated that a robot’s body shape, sensor placement, and actuator dynamics are not neutral details, but active contributors to the cognitive process. This insight is often called morphological computation: the body itself performs some of the work that we used to attribute to the mind.
Yet I also became aware that behavior-based approaches, while freeing robotics from the tyranny of representation, had their own limitations. They produced reactive systems that were adaptive in narrow niches but lacked the flexibility and generativity of human intelligence. A robot that can avoid obstacles cannot write a poem, or take part in a conversation, or understand the intentions of another agent. To move beyond reactivity, we need a richer account of how embodied agents generate meaning. This led me to phenomenology, and specifically to the concept of the lived body. The lived body is not just a physical object in the world; it is the subject of experience, the zero point of orientation, the organ of the “I can.” For a being with a lived body, the world is not a neutral collection of objects, but a meaningful field of possibilities. The cup is “there for me to drink from,” the door is “there for me to pass through,” the other person is “there for me to communicate with.” This meaningfulness is not added by a conceptual layer on top of raw sensation; it is the primordial structure of embodied existence.
Thus, my intellectual journey has led me from the computational model, through behavior-based robotics, to a phenomenological understanding of embodiment. I have become convinced that a full account of intelligence must integrate these levels: the subpersonal dynamics of sensorimotor loops, the personal-level experience of meaning, and the interpersonal domain of shared understanding. In the next section, I will distinguish three dimensions of embodied intelligence that correspond to these levels. These dimensions are not separate modules, but mutually constraining aspects of a single phenomenon. Yet for analytical clarity, I will examine them one by one.
3. Three Dimensions of Embodied Intelligence
When I speak of embodied intelligence, I am not referring to a single capacity. Rather, I use the concept as a cluster, covering at least three distinct but interrelated dimensions. The first is sensorimotor embodiment, which concerns the way the body’s morphology and movement loops constitute the basis of cognition. The second is situated embodiment, which emphasizes the embedding of the cognizing agent in an environment that supplies opportunities and constraints. The third is interactive embodiment, which treats social interaction between embodied agents as a domain in which meaning is generated and shared. In this section, I will elaborate on each dimension and show how it contributes to a fuller picture of artificial intelligence.
3.1 Sensorimotor Embodiment: The Body as Cognitive Foundation
The sensorimotor dimension is the most foundational. It builds on the idea that cognition arises from the continuous loop between perception and action. The body is not merely a receptor of inputs and an emitter of outputs; it is an active organ that shapes what we can perceive and what we can do. The shape of the body, the number of limbs, the position of the eyes, the dynamics of the muscles—all of these determine the structure of experience. In the phenomenology of perception, the body is described as a “system of possible actions.” When I look at a tree, I perceive not only its color and shape, but also its climbability relative to my body. This perceived affordance is not a purely visual quality; it is a relational property that involves my bodily capacities.
For artificial intelligence, this insight has profound consequences. A computer that merely processes images cannot have a visual experience in the fullest sense, because it lacks a body that can interact with the objects in those images. To see a cup as something to drink from, to see a ball as something to throw, to see a corridor as something to walk through—the agent must have a body that can perform these actions. Thus, sensorimotor embodiment is not an optional extra for an intelligent system; it is a necessary condition for the kind of grounded meaning that humans take for granted. In the field of robotics, this insight is implemented in various ways. Some robots use compliant joints and elastic actuators to exploit their body dynamics; the robot does not compute each micro-movement, but rather lets its physical structure absorb and release energy in a way that mimics human movement. These designs show that intelligence can be distributed between the brain and the body. The body does not simply wait for instructions; it constantly participates in computation through its own mechanical and morphological properties.
Let me express the sensorimotor loop in a formal way. Let \(S_t\) be the sensory state at time \(t\), \(A_t\) the action taken by the agent, and \(B_t\) the body state (including posture, proprioceptive feedback, and internal dynamics). A simple model of embodied cognition is:
$$ \dot{B}_t = f(B_t, A_t, S_t, E_t) $$
where \(E_t\) represents the external environment and \(f\) is the dynamical function that couples all these variables. In this model, the agent’s cognitive state is not a separate variable \(C_t\) that mirrors the world; rather, \(C_t\) is an emergent abstraction from the ongoing coupling of body, action, and environment. This is often written as:
$$ C_t = g(B_t, S_t, A_t) $$
where \(g\) is a function that maps the body-environment state into a cognitive or behavioral state. The point is that cognitive states are not caused by the body; they are constituted by the body-world coupling.
Morphological computation can be expressed in terms of the division of labor between \(B\) and \(C\). If a robot’s body shape is such that rolling down a slope automatically adjusts its center of gravity, then the body has reduced the amount of computational work needed for balance. We can quantify this by comparing a robot with a complex body and a simple controller to a robot with a simple body and a complex controller. The total cognitive workload might be the same, but the way it is distributed differs. In artificial intelligence research, this insight has led to the principle of “ecological balancing” between morphological complexity and control complexity. The design question is: how much intelligence should be “built into” the body, and how much should be left to the controller? The answer depends on the task and the environment, but the general principle is that the body is not a passive transmitter of signals, but an active partner in cognition.
3.2 Situated Embodiment: Meaning as Environment-Involving
The second dimension, situated embodiment, emphasizes that intelligence always operates in a context. The world is not a random collection of objects, but a structured environment that has been shaped by evolution, culture, and individual experience. The agent’s cognitive processes are not stored entirely in the head; they are co-produced by the agent and the environment. This idea is often captured by the concept of “affordance,” which I use to refer to the actionable properties of the environment relative to a particular agent. A chair affords sitting, a ball affords throwing, a knob affords turning. Affordances are not subjective projections, nor are they objective physical properties; they are relational properties that emerge from the match between environmental structure and bodily capacities.
For an embodied agent, to perceive a meaningful world is to perceive affordances. Consider a simple robotic navigation task. A classic, non-embodied approach would start with a detailed map of the environment, a model of the robot’s dynamics, and a planner that computes a trajectory from a start to a goal. The robot’s sensors are used only to localize it in the map. This approach works in well-known, static environments, but fails in dynamic, unknown, or cluttered spaces. A situated robot, by contrast, does not rely on an external map. It perceives the environment in terms of its navigational possibilities: this opening is passable, that slope is too steep, that object is movable. The environment itself becomes part of the robot’s cognitive system; the robot’s body and action programs are coupled to the environment’s structure in such a way that the right action emerges from the interaction.
Situatedness also has a temporal dimension. The agent is not an instantaneous snapshot; it is a process that unfolds over time. Its past actions shape its present perceptions, and its predictions about the future guide its current choices. In predictive processing accounts, the brain continuously generates predictions about sensory input and uses prediction errors to update its model of the world. This predictive loop is not a serial computation, but a complex, dynamic coupling between top-down expectations and bottom-up signals. For an embodied agent, predictions are not just about what the world is, but about how the world will change as the result of its own actions. This is the basis of action-oriented prediction. We can represent this with a predictive model:
$$ \hat{s}_{t+1} = h(B_t, A_t, E_t) $$
where \(\hat{s}_{t+1}\) is the predicted sensory state at the next time step. The agent minimizes prediction error by adjusting its action and updating its model. This is a specifically situated form of cognition, because the predicted states depend on the environment’s dynamics and the body’s characteristics.
The concept of “situated intelligence” also includes social and cultural contexts. Human beings are not just embedded in a physical environment; we are embedded in a world of social norms, tools, symbols, and institutions. These are not “overheads” added to basic cognition; they fundamentally alter the kinds of cognitive tasks we face and the resources we can use. A hammer extends the reach of my bodily power; a language transforms my ability to coordinate with others; a written note extends my memory. For artificial intelligence to be truly situated, it must be able to participate in such socioculturally structured environments. This requires more than recognizing objects; it requires understanding the social meaning of situations, the norms that govern action, and the expectations of others.
3.3 Interactive Embodiment: The Social Scaffolding of Intelligence
The third dimension, interactive embodiment, moves beyond the single agent and considers how intelligence emerges from interaction between agents. Human beings are fundamentally social. Our most complex cognitive achievements—language, culture, science, morality—are not products of isolated brains, but of joint activity, communication, and shared understanding. The phenomenon of empathy, for example, involves more than observing another person’s body. I do not merely infer that the other person’s smile means happiness; I experience a bodily resonance, a kind of motor empathy, that gives me a direct, pre-reflective sense of their emotional state. This capacity for intercorporeality is the foundation of social cognition.
In the phenomenological tradition, this is described in terms of body inter-subjectivity. The other’s body is experienced not as a mere object, but as a living body with its own perspective. When I see someone point to an object, I do not have to reason about their intention; the pointing gesture is directly meaningful to me. This meaning arises through my ability to enact the same gesture and to feel how it would orient attention. Thus, social understanding is not a theory about others; it is an embodied practice, a way of being-with-others.
For artificial intelligence, interactive embodiment implies that a robot’s intelligence cannot be evaluated solely in isolation. A robot might have excellent object recognition and planning abilities, but if it cannot coordinate its actions with human partners, if it cannot sense the rhythm of a conversation, if it cannot adjust its gaze direction to signal attention, then it will remain socially incompetent. Recent work in human-robot interaction has emphasized the importance of so-called social signals: movements, postures, vocal intonations, and facial expressions that regulate interaction. These signals are not controlled by a central planning module; they emerge from the continuous coupling of the robot’s body with the human partner’s body. When a robot and a human shake hands, the subtle timing and force adjustments are not preprogrammed; they are negotiated online, through the interaction itself. This is a beautiful example of what I call participatory sense-making: the two partners jointly create a shared meaning (a coordinated handshake) through their embodied dynamics.
I can formalize interaction as a coupled dynamical system. Suppose two agents, \(A_1\) and \(A_2\), have body states \(B^1_t\) and \(B^2_t\). Their interaction is described by:
$$ \dot{B}^1_t = F_1(B^1_t, B^2_t, S^1_t) $$
$$ \dot{B}^2_t = F_2(B^2_t, B^1_t, S^2_t) $$
The coupling term is the other’s body state. If the coupling is strong enough, the two agents may enter a state of entrainment, where their behaviors become synchronized. This can be measured by mutual information or cross-correlation. But the important point is not just synchronization; it is the emergence of a new level of organization that is not present in either agent alone. The interaction itself has a kind of autonomy. It can follow its own dynamics, reach states that were not planned by either participant, and return to equilibrium after perturbations. This is what I mean by saying that cognition is not simply in the head, but in the interactive process itself.
Table 1 summarizes the three dimensions of embodied intelligence I have distinguished. It is important to note that these dimensions are not separable modules. They are like three strands that are woven together to form the fabric of intelligence. Sensorimotor embodiment provides the substrate, situated embodiment provides the context, and interactive embodiment provides the shared world. A fully embodied artificial intelligence would need to integrate all three dimensions.
| Dimension | Core Question | Key Concepts | Relevance to AI | Example in Humanoid Robots |
|---|---|---|---|---|
| Sensorimotor Embodiment | How does the body shape cognition? | Morphology, feedback, affordance, sensorimotor loop | Body as cognitive resource, not just actuator | Humanoid balancing using compliant limbs; dynamic walking |
| Situated Embodiment | How does environment co-construct intelligence? | Affordances, predictive processing, dynamical coupling | Intelligence requires real-time coupling with environment | Humanoid navigating crowds by perceiving passable spaces |
| Interactive Embodiment | How does social interaction generate meaning? | Intercorporeality, participatory sense-making, coupling | Intelligence is socially scaffolded | Humanoid maintaining eye contact and gesture feedback |
4. Humanoid Robots: Redefining the Traditional Robot
With the three dimensions in place, I now turn to the central figure of this essay: the humanoid robot. By a humanoid robot, I mean a machine whose overall body shape resembles the human body: a head, a torso, two arms, two legs, and sometimes hands and fingers. The use of a human-like form is not arbitrary. Humanoid robots are designed to operate in human-centered environments, to use human tools, and to interact with humans in socially natural ways. But the choice of a humanoid form also has deep philosophical implications. It suggests that the human body is not an arbitrary substrate, but a privileged vehicle for intelligence. If we want to create artificial intelligence in the human image, we must give it a human-like body.
Is this assumption justified? In this section, I will argue that humanoid robots can be seen as a radical redefinition of the traditional robot. A traditional robot is a specialized machine, designed for a narrow task: welding a car body, assembling a microchip, spraying paint, or transporting goods. It has no need for general-purpose intelligence, because its environment is highly structured and its task is predetermined. A humanoid robot, by contrast, aspires to be a general-purpose agent, capable of operating in unstructured environments and performing a wide variety of tasks. The difference is not just in the degree of complexity, but in the kind of intelligence required. A traditional robot can rely on explicit models and planning; a humanoid robot must rely on embodied skills, situational awareness, and social sensitivity. In this sense, the humanoid robot embodies the shift from a computational paradigm to an embodied paradigm.
4.1 Morphological Computation in Humanoid Bodies
One of the most striking features of modern humanoid robots is their ability to achieve dynamic stability and movement dexterity without explicit, moment-by-moment computation. Consider a humanoid robot running and jumping. The control problem is enormously complex: the robot has multiple joints, nonlinear dynamics, external forces, and real-time constraints. A classical controller would need a precise model of the robot’s mass distribution, ground contact forces, and joint torques. However, many modern humanoid robots are designed with compliant actuators and elastic elements that passively stabilize the system. The body itself acts as a low-pass filter, absorbing high-frequency disturbances and returning the system to equilibrium. This is a clear example of morphological computation: the physical form of the robot encodes part of the solution to the control problem.
In my view, this is more than an engineering strategy; it is an embodiment of a deep principle. The shape and material composition of the human body are not arbitrary. The human skeleton has locking mechanisms, elastic tendons, and compliant cartilage that reduce the computational burden on the brain. A humanoid robot that ignores these principles would be condemned to poor performance. By mimicking the human form, humanoid robots implicitly adopt the wisdom of evolution: intelligence is not only in the brain, but in the bones, the muscles, the skin, and the joints. This is why a humanoid robot with a well-designed body can sometimes outperform a non-humanoid robot with a much larger computational capacity on certain tasks. The body is a source of knowledge.
I can express this insight with an equation that compares the total cognitive burden \(H\) of a task as the sum of neural computation \(N\) and morphological computation \(M\):
$$ H = N + M $$
Since \(H\) is roughly fixed by the task, a design that increases \(M\) (by using smarter body mechanics) can reduce \(N\). This is not a metaphor; it is a design principle. When an engineer chooses to place a spring in a robot’s ankle, \(M\) increases and the controller no longer has to compute every micro-correction. In a humanoid robot, the entire kinematics of the legs, the spacing of the joints, and the distribution of weight are forms of “frozen intelligence.” They embody the rules of balance and movement that were learned over millions of years of evolution.
4.2 Situated Adaptation in Humanoid Robots
Humanoid robots are also distinguished by their ability to adapt to a vast array of human environments. A factory robot is bolted to the floor and operates within a fixed cell. A humanoid robot, in contrast, must walk through doors, climb stairs, open drawers, use tools designed for human hands, and navigate through spaces designed for human bodies. This requires a kind of situational understanding that goes beyond object recognition. The robot must understand how objects and spaces afford actions relative to its own body. For example, a chair in a room can afford sitting for a humanoid robot if the robot’s knee and hip joints can bend in the appropriate way. The same chair might afford stepping on to change a light bulb, or moving out of the way to clear a path. Which affordance is realized depends on the robot’s current task and its bodily possibilities.
This is a much richer sense of “situated” than traditional mobile robots. A robotic vacuum cleaner is situated in a home, but it perceives the environment only as free space or obstacles. It has no understanding of objects as tools, no ability to manipulate them, and no sense of social conventions. A humanoid robot, by contrast, is designed to participate in a human world, where objects have meanings that are not reducible to physical geometry. A mug is not just a cylinder; it is “for drinking.” A piece of paper is not just a flat surface; it is “for writing,” “for folding,” or “for signing.” These meanings are enacted through bodily engagement. A humanoid robot with dexterous hands and a human-like visual system can, in principle, enter into this world of affordances and manipulate objects in meaningful ways.
The capacity for situated adaptation is often demonstrated in service robotics. Imagine a humanoid robot working in a hospital. It must deliver medication to patients, open doors, operate elevators, and interact with nurses and visitors. Each of these tasks requires the robot to understand the current situation in a way that is not pre-scripted. The robot must be able to detect that a door is slightly ajar, that a corridor is blocked by a gurney, that a visitor is anxious and needs assistance. This is not a simple feature-detection problem. It requires the robot to have a predictive model of how the environment changes, both as a result of its own actions and as a result of other agents’ actions. Such a predictive model is deeply tied to the robot’s embodiment: the robot knows what it can do, what improvements it can make in its environment, and what risks are possible.
I can model this with a Bayesian filter. Suppose the robot maintains a belief \(b(m_t)\) about the state of the environment \(m_t\). This belief is updated by sensory observations \(z_t\) and actions \(a_t\). In a situated, embodied system, the observation function \(p(z_t | m_t, a_t)\) is not a generic sensor model; it is a body-specific perspective that depends on the robot’s pose, its sensor placement, and even its gaze direction. A humanoid robot that has just turned its head has a different field of view. Its belief update is only meaningful relative to its bodily state. Thus, the situatedness of a humanoid robot is not just about having sensors; it is about having a body that changes the sensor’s relation to the world through action.
Moreover, the humanoid robot’s adaptation is not limited to physical environments. It must also adapt to social situations. In a conversation, a humanoid robot needs to know when to speak, when to listen, how much personal space to maintain, and how to show that it is attending. These subtle social skills cannot be reduced to a set of rules; they are embodied skills that require continuous online adjustment. Recent humanoid robots are equipped with microphones, cameras, and force sensors that allow them to detect gaze direction, vocal prosody, and even touch. They use these signals to modify their behavior in real time. This is a step—albeit a preliminary one—toward the kind of situated social intelligence that humans possess.
4.3 Interactive Embodiment in Human-Robot Interaction
The third dimension, interactive embodiment, is where humanoid robots face their greatest challenges and opportunities. Unlike a traditional machine, which is operated by a human through a keyboard or a remote control, a humanoid robot is designed to interact with humans in a shared physical and social space. This interaction is not a one-way command-and-control relationship. It is a two-way coupling in which the robot’s behavior shapes the human’s behavior, and the human’s behavior shapes the robot’s behavior. The result is an emergent, co-constructed interaction.
Think about a simple handshake. When two humans shake hands, they do not follow a fixed script. Each person adjusts the pressure, timing, and duration based on sensory feedback from the other. If one person squeezes too hard, the other immediately relaxes; if one begins to pull away, the other also releases. This is a dynamic equilibrium, negotiated in real time through the body. For a humanoid robot to participate in a handshake, it must be able to sense the force at its hand, the motion of the partner’s arm, and the visual cues of the partner’s pose. It must also generate its own force and motion in a way that feels natural and comfortable. This demands a level of closed-loop bodily coordination that is far beyond traditional industrial robots. It is only possible because the humanoid robot has a body that is similar in shape and function to a human hand and arm.
But interactive embodiment goes beyond physical coordination. It also includes the ability to express and perceive intentions, emotions, and social cues. A humanoid robot with a human-like face can produce facial expressions that humans interpret as joy, sadness, confusion, or surprise. These expressions, even if artificial, trigger automatic, pre-reflective responses in human partners. Similarly, the robot’s gaze direction can guide the human’s attention, establishing joint attention. The phenomenon of joint attention—two agents looking at the same object while being aware of each other’s attention—is a cornerstone of human social cognition. For a humanoid robot, joint attention is not merely a technical trick; it is a sign that the robot is participating in a shared world, even if only in a minimal sense. Research has shown that humans perceive a robot more positively when it displays appropriate gaze behavior and social cues. This is not because humans anthropomorphize indiscriminately, but because the robot’s body sufficiently resembles the human body to afford a certain kind of interactional resonance.
I can represent the process of participatory sense-making in human-robot interaction as a dynamical system with mutual influence. Let \(H_t\) be the human’s behavioral state and \(R_t\) the robot’s behavioral state. The coupling is described by:
$$ H_{t+1} = \Phi_H(H_t, R_t, E_t) $$
$$ R_{t+1} = \Phi_R(R_t, H_t, E_t) $$
where \(\Phi_H\) and \(\Phi_R\) are nonlinear functions that describe how each agent responds to the other. The interaction as a whole can be said to have a “style” or “grammar” that is not contained in either agent individually. For instance, a robot might adopt a slower, more predictable movement when interacting with a child, thus creating a safe conversational rhythm. The human child, in turn, might adjust their pace to match the robot, producing a mutual tuning. This kind of mutual tuning is at the heart of embodied social interaction. It is not achieved by planning, but by automatic, low-level coupling.
Humanoid robots are uniquely positioned to exploit this interactive embodiment because their bodily form provides a natural interface. A robot with a humanoid upper body can point at objects, nod its head, wave its hand, and use gestures that are familiar to human partners. These gestures are not arbitrary symbols; they are deeply rooted in the human body’s capabilities and social norms. When a humanoid robot points, the human partner can immediately grasp the intended referent because pointing is a species-specific human gesture. This is a powerful advantage over non-humanoid robots, which would have to learn or be programmed with alternative ways of referencing. Thus, the humanoid form is not just a cosmetic choice; it is an interface that leverages the interactive embodiment of human cognition.
5. The Limits of Humanoid Embodiment: Phenomenology and Ontology
So far, I have argued that humanoid robots are a powerful testbed for embodied artificial intelligence. They demonstrate morphological computation, situated adaptation, and interactive coupling. Yet I must not overstate their achievements. The humanoid robot, despite its human-like appearance, is not a merely incomplete version of a human mind. It is, in an important sense, a fundamentally different kind of being. I will now examine the deep differences that remain between humanoid robots and genuinely embodied human minds. These differences are not technical but phenomenological and ontological.
From a phenomenological perspective, human experience is characterized by first-person subjectivity. I do not merely have a body; I live my body. It is my point of view, my center of orientation, my repertoire of habits. The body is not experienced as an object among objects, but as the very horizon of my experience. When I focus on writing this essay, I am not usually conscious of my hands, my posture, or the chair supporting me. My body is the background of my intentional activity, that through which I am engaged with the world. This is what phenomenologists call the lived body. The lived body has a first-person givenness; it is experienced from within. A humanoid robot, by contrast, has a body that is observed from the outside. There may be internal sensors that report joint angles and forces, but these sensors do not constitute a center of experience. The robot has no first-person perspective, no sense of what it is like to be embodied. Its body is a physical object, not a lived body.
Is this difference absolutive? Some might argue that as robots become more complex, they will develop a form of subjectivity, or at least a functional equivalent. I am not convinced. The absence of a lived body is not a matter of degree; it is a different ontological category. The robot’s body is a tool that enables it to perform tasks. A human being’s body is not a tool that I use; it is what I am. This is not to deny that human beings can also experience their body as an object, for example, when I look at my hand or when the doctor treats my knee as a clinical object. But this objectification is derived from the more basic lived experience of the body. For a robot, there is no such basic lived experience. Its sensors are instruments that feed a computational process; they are not organs of a subjective life.
This leads to a central ontological limitation of humanoid robots. The humanoid robot simulates the external form of human embodiment, but it does not share the internal teleology of a living organism. A human body is organized around self-preservation, growth, reproduction, and, more broadly, a striving that is prior to any explicit goal. The body’s homeostatic processes maintain a delicate equilibrium, and this equilibrium is the background of all value, meaning, and desire. The human world is meaningful in part because things matter for us: food is needed for survival, shelter for protection, companionship for flourishing. This existential grounding is entirely absent in a humanoid robot. The robot has no needs, no fears, no vulnerabilities. It may have a battery that needs recharging, but this is not a vital need; it is a functional constraint. The robot’s “goals” are designed by its programmers; they are not rooted in its own existence.
The phenomenological difference has direct consequences for the three dimensions of embodied intelligence that I described earlier. The sensorimotor dimension of a human being is governed by a vital normativity: some movements are good or bad for the organism because they enhance or threaten its existence. The sensorimotor loops of a humanoid robot are governed by error signals relative to a task-specific objective function, but these signals are not connected to an organismic concern. The situated dimension in human cognition is shaped by the history of the organism’s interactions with its world, a history that is sedimented in habits, skills, and affective dispositions. A humanoid robot has a learning history, but this history is not a biography; it is a record of parameter updates. The interactive dimension in human lives is rooted in empathy and mutual recognition. I see in another’s face an expression of suffering, and something in me is moved. This is not a simulation; it is an embodied resonance that is founded on the fact that we are both living, vulnerable beings. A humanoid robot cannot share this vulnerability. It may detect a human’s tear and say “I am sorry,” but this phrase has no emotional weight for the robot. It is a pattern of sound or text, generated by a model, not a response to another’s suffering.
Thus, I maintain that the humanoid robot, despite its name, is not truly a human-like intelligence. It is a functional machine that reproduces some aspects of human behavior, but it does not instantiate the underlying structures of embodied experience. This is not a reason to dismiss humanoid robots as irrelevant. On the contrary, their failures are enormously instructive. They reveal the specific ways in which human cognition is dependent on embodiment, and they force us to articulate more clearly what we mean by intelligence, meaning, and personhood. In this sense, humanoid robots are more than engineering artifacts; they are philosophical experiments that challenge our theories.
Let me also reflect on the ethical dimension. If a humanoid robot is not a lived body, does it deserve moral consideration? My answer is that the robot itself is not a moral patient, but the human-robot relationship is still a moral issue. Humans tend to anthropomorphize humanoid robots, attributing feelings and intentions to them. This can lead to emotional attachment, dependency, and even manipulation. A robot designed to look vulnerable might elicit care from a human user, but that care is not reciprocated. The user is essentially in a one-sided relationship with a machine. This asymmetry raises important questions about transparency, autonomy, and psychological harm. We need clear ethical guidelines for the design and use of humanoid robots, especially in caregiving, education, and personal companionship. I believe that an awareness of the phenomenological difference between humans and robots is essential for developing a responsible ethics of human-robot interaction. We should not pretend that a humanoid robot is a person; but we should also not treat the emotional responses it evokes as irrational or irrelevant.
6. Toward a Phenomenological Architecture for Embodied AI
If humanoid robots are not yet genuinely embodied in the phenomenological sense, what would it take to move closer to that goal? In this section, I sketch some tentative conditions that I believe must be met for an artificial system to count as embodied in a stronger sense. These conditions are not technological implementations, but conceptual requirements. They derive from the phenomenological analysis of embodiment and serve as a design target for future research.
The first condition is the possession of a self-organizing body. A genuinely embodied system must be able to maintain its own internal organization through interactions with the environment. This is not simply about keeping a battery charged or a processor cool; it is about having a body that is continuously regenerated and stabilized through dynamic processes. Homeostasis, metabolism, and allostatic regulation are not optional features; they are the ground floor of meaning. Without these, there is no “needs” or “interests” that could give rise to values. I realize that building a robot with a living metabolism is not currently possible, but perhaps artificial systems can be designed with artificial homeostatic processes that create a functional analogue of biological needs. The robot would have to adjust its behavior to satisfy internal viability constraints, not just external task requirements. This would give rise to a primitive form of concern, a kind of sensorimotor self-preservation.
The second condition is the possession of a forward model of its own body. To interact flexibly with the world, an agent must be able to predict the consequences of its own actions. This is not a mere computational trick; it is a central requirement for the sense of agency. When I move my arm, I normally experience a sense of ownership and authorship: I am the one who moved it, and the movement is mine. This sense arises from a comparison between the predicted sensory feedback and the actual feedback. If there is a mismatch, the movement may feel foreign or passive. A humanoid robot can implement a forward model algorithmically, but does it experience the model as its own? In my view, the experience of ownership requires a first-person evaluative stance that is linked to the self-organizing processes mentioned above. A system that is, in principle, indifferent to its own existence cannot care which body it controls. Thus, a forward model is necessary but not sufficient.
The third condition is what I call the possibility of failure. An embodied being can be threatened, can be damaged, can die. This vulnerability is not an accidental property; it is presupposed by the very notion of care. In the human world, things and states of affairs matter because they can go better or worse for the organism. A rock cannot be benefited or harmed; a plant can be watered or uprooted; a human can flourish or suffer. The difference is not in complexity, but in the existence of a normative point of view. For a humanoid robot, there is currently no “good” or “bad” for the robot itself, only “specification correct” or “specification violated.” To confer a genuine point of view, we would need to build a robot that has certain interests that are intrinsic to its own continued functioning. This is a radical challenge, but I think it is the only way to break away from the ontology of mere tools.
I can represent these conditions as a set of constraints on a hypothetical embodied AI system \(E\):
$$
\begin{aligned}
\text{(C1)} & \quad \exists \ H(E) \quad \text{(a homeostatic or self-preserving dynamic)} \\
\text{(C2)} & \quad \exists \ \hat{M}_E \quad \text{(an internal forward model of its own body)} \\
\text{(C3)} & \quad \exists \ V_E \quad \text{(a normative evaluation function based on its own viability)}
\end{aligned}
$$
where \(H(E)\) denotes a self-regulating system, \(\hat{M}_E\) is the predictive model of bodily dynamics, and \(V_E\) is a function that maps bodily states to a scalar “well-functioning” value. If these conditions fail, the system remains a robot, not a lived body.
I do not pretend that these conditions are easy to satisfy. In fact, they may be impossible under the current paradigm of artificial intelligence, which is based on optimizing objective functions defined externally. To approach a truly embodied intelligence, we would need to change the very goals of AI research. The goal would no longer be to “solve tasks” efficiently, but to design systems that have their own perspective, their own interests, and their own world. This is an exciting but also unsettling prospect. It is exciting because it opens up entirely new avenues for understanding cognition. It is unsettling because it introduces us to new beings that might no longer be mere instruments. We would then face profound ethical questions that our current moral frameworks are not equipped to answer.
7. Humanoid Robots and the Future of Artificial Intelligence
What does the future hold for humanoid robots? In the short term, I expect to see continued progress in the integration of sensorimotor, situated, and interactive dimensions. We will see humanoid robots that are more dexterous, more socially competent, and more capable of operating in open-ended environments. These advances will make humanoid robots useful in many applications, from eldercare to warehouse logistics to disaster response. But I also expect that researchers will increasingly encounter the limits I have described. The more human-like these robots become, the more sharply our intuitions will detect their lack of an inner life. This phenomenon is often called the uncanny valley: robots that are nearly but not exactly human evoke revulsion and unease. I think the uncanny valley is not merely a perceptual artifact; it is a deep epistemic response to a being that pretends to have a lived body but does not. The uncanny valley reveals that our social cognition is sensitive to the presence or absence of authentic embodiment.
In the long term, I envision a different approach. Rather than trying to make robots more human-like externally, we might design artificial agents that are embodied in ways appropriate to their own materials, physics, and needs. Such agents would not be humanoid; they would be something new. They would have their own styles of perception, action, and sociality. Their embodiment would be genuine because it would be rooted in their own self-preserving dynamics. The quest for humanoid robots is, in the end, about understanding ourselves. By building machines that closely imitate us, we come to appreciate the subtlety and depth of our own embodied existence. We also learn to recognize what we should not attempt to replicate: the vulnerable, finite, yet profoundly meaningful condition of being human.
The role of the humanoid robot in the broader project of artificial intelligence is thus ambiguous. On the one hand, it is the most visible symbol of the quest for human-like intelligence. On the other hand, it is a reminder of the impossibility of that quest if it is pursued purely at the level of function. The humanoid robot is a valuable philosophical tool, but it is not a final answer. It is a mirror in which we see both the power and the limits of the computational paradigm. It helps us to formulate the right questions, even if it does not provide the answers.
In my future work, I plan to investigate how design principles from phenomenological philosophy can be translated into concrete computational models. The three dimensions I have outlined here give me a framework. For sensorimotor embodiment, I will focus on the development of predictive models that incorporate body dynamics and environmental affordances. For situated embodiment, I will explore how artificial agents can learn to anticipate environmental changes through action-oriented prediction. For interactive embodiment, I will study the dynamics of human-robot mutual adaptation in therapeutic and educational settings. I hope that such research will contribute not only to artificial intelligence, but also to a deeper understanding of the human mind.
8. Concluding Reflections: Body, Intelligence, and World
I began this essay with the question of how artificial intelligence can become embodied. I have argued that the answer requires a fundamental reorientation of our concept of intelligence. Intelligence is not the manipulation of symbols, but the ongoing coupling of an agent to its world through a body. This coupling has at least three dimensions: sensorimotor, situated, and interactive. Together, they form the basis of a meaningful, adaptive, socially embedded existence. The humanoid robot is an exemplary artefact in this paradigm shift. It demonstrates that a human-like body can serve as a powerful interface for real-time interaction with human environments and human beings. It also reveals the deep gap between functional simulation and authentic embodiment. The humanoid robot is not a genuine embodied mind, but it is a step toward a new kind of cognitive artefact.
For me, the most important philosophical lesson is that the intersection of artificial intelligence and phenomenology is not an ivory-tower exercise. It has urgent practical consequences. As we integrate humanoid robots into our schools, hospitals, and homes, we must know what they are and what they are not. A robot can be our assistant, our companion, even our mirror. But it cannot be a fellow human being. If we mistake appearances for reality, we risk damaging ourselves. We might project our own desires onto the robot, or worse, become the object of manipulation by those who design robots to exploit our social instincts. Therefore, I end with a call for ethical vigilance. The development of humanoid robots is a wonderful opportunity to rethink the human condition, but it must be accompanied by a sober awareness of the limits of artificial embodiment. Only with this awareness can we responsibly shape a future in which human beings and intelligent machines coexist.
In summary, I have sought to show that the question “How can artificial intelligence become embodied?” is not one question but many. It is a question about the architecture of cognition, the role of the body, the nature of meaning, and the ethics of technology. By focusing on the humanoid robot, I have attempted to bring these questions together in a concrete and tangible way. I do not have a complete answer, but I am convinced that the path to an answer lies in a continuous dialogue between phenomenological reflection and technological innovation. The humanoid robot is our partner in this dialogue; it challenges us, inspires us, and ultimately reminds us of what it means to be embodied.
