Humanoid robots have progressively expanded their social roles, entering domains such as customer service, healthcare, and education. One of the most influential aspects of humanoid robot design is the facial appearance, which directly shapes users’ emotional reactions and willingness to interact. Although many studies have examined individual facial attributes of humanoid robots, the majority rely on subjective self-reports. In this work, we aimed to provide a more objective evaluation of how specific facial features of humanoid robots affect users’ affective perception. To this end, we combined subjective ratings, eye-tracking, and electroencephalographic (EEG) measurements to investigate the effects of facial width-to-height ratio, face shape, eye shape, and mouth expression on user emotional cognition.
Prior evidence suggests that users pay more attention to the face of a humanoid robot than to other body parts. The face contains key social signals, and subtle changes in geometry or expression can dramatically alter perceived warmth, competence, and appeal. In the present study, we systematically manipulated four facial features of humanoid robot prototypes: facial width-to-height ratio (high, medium, low), face shape (round, square), eye shape (round, square), and mouth expression (smiling, neutral linear). Our central research question was which of these features dominate users’ affective evaluation and early visual and neural processing. We hypothesized that different features would produce distinct effects on subjective liking, visual attention allocation, and event-related potential (ERP) components such as N1, P2, and P3.
To test these hypotheses, we first created twenty-four humanoid robot face prototypes by combining all levels of the four feature dimensions. We then conducted a screening session to select the most preferred facial width-to-height ratio. After confirming that a medium ratio yielded the highest liking ratings, we retained eight prototypes for the main eye-tracking and EEG experiments. Each retained humanoid robot face had a medium width-to-height ratio and varied orthogonally in face shape (round vs. square), eye shape (round vs. square), and mouth expression (smiling vs. neutral linear). All images were standardized at 670×520 pixels against a white background. A sample of the humanoid robot face prototype is illustrated below.

Methods
Participants
Twenty-two undergraduate students (eleven males, eleven females) participated in the experiment. Their ages ranged from 18 to 22 years (mean age ± standard deviation = 20.3 ± 1.2 years). All participants were right-handed, reported normal or corrected-to-normal vision, and had no history of neurological or psychological disorders. Prior to the experiment, each participant read a detailed information sheet and signed an informed consent form. The experimental protocol was approved by the local ethics committee. After the session, each participant received a small gift as compensation.
Experimental Materials
We constructed the humanoid robot face stimuli using professional three-dimensional modeling software, following design conventions observed in popular commercial humanoid robots and previous research. The base model included a human-like pupil, realistic ears, and a linear or curved mouth. We systematically varied four attributes: facial width-to-height ratio (high, medium, low), face shape (round, square), eye shape (round, square), and mouth expression (smiling, neutral linear). All combinations yielded twenty-four unique images. In a preliminary subjective evaluation, participants rated their liking for the three width-to-height ratio levels on a seven-point scale from −3 (“disliked”) to +3 (“liked”). A repeated-measures ANOVA revealed a significant main effect of width-to-height ratio, F(2, 42) = 6.347, p = 0.002, partial η² = 0.025. Post-hoc comparisons indicated that the medium ratio (M = 0.339, SD = 1.757) was significantly preferred over both the low ratio (M = −0.185, SD = 1.673, p = 0.006) and the high ratio (M = −0.292, SD = 1.780, p = 0.001). Therefore, we used the eight prototypes with the medium width-to-height ratio as the formal experimental stimuli.
EEG Procedure
The EEG session was conducted in a quiet, temperature-controlled laboratory with dimmed lighting. We used a Neuroscan 64-channel SynAmps system (Compumedics, USA) to record continuous EEG. Participants sat comfortably in front of a 22-inch monitor while the display sequence was controlled by a computer. Each trial began with a fixation cross on an otherwise blank screen presented for 800–1000 ms. Then the humanoid robot face stimulus appeared for 800–1000 ms, and participants were asked to freely view the image. Following the stimulus offset, a response screen requested participants to indicate whether they liked the humanoid robot face by clicking the left mouse button (like) or the right mouse button (dislike). The order of the eight stimuli was randomized across participants, and each stimulus was presented multiple times to accumulate enough trials for ERP averaging.
Eye-Tracking Procedure
The eye-tracking session took place in a human factors laboratory with controlled ambient light and acoustic noise. We used the SMI iView RED remote eye-tracking system (SensoMotoric Instruments, USA) with a sampling rate of 500 Hz. Each participant was seated approximately 70 cm from a 22-inch display. We performed a nine-point calibration procedure, requiring the average deviation for both eyes to be less than 0.5°. In the formal test, participants viewed the eight humanoid robot face images in a natural, continuous sequence. Each image was displayed for 10 seconds, with a 1000–2000 ms blank interval between consecutive images. Gaze data were recorded for offline analysis.
Data Processing
For eye-tracking data, we used Begaze 3.6 software (SensoMotoric Instruments, USA) to define four areas of interest (AOIs): mouth, eyes, face, and background. These AOIs covered the principal facial regions of the humanoid robot. Two dependent measures were extracted: fixation count and fixation duration. Fixation count reflects the number of times a participant’s gaze landed within an AOI, while fixation duration represents the total dwell time in that AOI. These metrics are established indicators of visual attention and cognitive processing load.
For EEG data, we processed the continuous recordings using MATLAB R2023b (MathWorks, USA) with the EEGLAB toolbox. The preprocessing pipeline included: (1) re-referencing to the average of all electrodes; (2) applying a band-pass filter of 0.5–30 Hz; (3) removing non-neural artifacts using independent component analysis (ICA); and (4) segmenting the data into epochs from 200 ms before stimulus onset to 1000 ms after stimulus onset. Baseline correction was applied using the pre-stimulus interval. We then averaged the epochs in each condition to obtain stimulus-locked ERP waveforms. Based on the scalp topographies, we selected a set of centro-parietal and parieto-occipital electrodes (Cz, Pz, POz, C4, FCz, PO4) for statistical analysis of the N1, P2, and P3 components. For each component, we computed the mean amplitude within a specific latency window (N1: 80–120 ms; P2: 150–250 ms; P3: 300–450 ms).
For all statistical analyses, we used repeated-measures analysis of variance (ANOVA) with within-subject factors. When sphericity was violated, we applied the Greenhouse-Geisser correction. Effect sizes are reported as partial eta-squared (ηp²). Post-hoc pairwise comparisons were corrected using the Bonferroni method. The significance threshold was set at p < 0.05.
For a typical repeated-measures ANOVA, the F-statistic is defined as the ratio of the mean square for the effect to the mean square for the error:
$$F = \frac{MS_{\text{effect}}}{MS_{\text{error}}}$$
where \(MS_{\text{effect}} = \frac{SS_{\text{effect}}}{df_{\text{effect}}}\) and \(MS_{\text{error}} = \frac{SS_{\text{error}}}{df_{\text{error}}}\). This formula was used to evaluate the significance of each facial feature on the dependent measures. Additionally, we computed descriptive statistics using the sample mean:
$$\bar{X} = \frac{1}{N}\sum_{i=1}^{N} X_i$$
and the sample standard deviation:
$$SD = \sqrt{\frac{1}{N-1}\sum_{i=1}^{N} (X_i – \bar{X})^2}$$
These formulas facilitated the summary of subjective, eye-tracking, and EEG data across conditions.
Results
Subjective Ratings
The subjective liking ratings were analyzed using a repeated-measures ANOVA with factors of face shape (round vs. square), eye shape (round vs. square), and mouth expression (smiling vs. neutral linear). The means and standard deviations for each condition are presented in Table 1.
| Face Shape | Eye Shape | Mouth Expression | Mean Liking | SD |
|---|---|---|---|---|
| Round | Round | Smiling | 0.772 | 1.239 |
| Round | Round | Neutral | 0.104 | 1.437 |
| Round | Square | Smiling | 0.680 | 1.310 |
| Round | Square | Neutral | 0.032 | 1.502 |
| Square | Round | Smiling | 1.208 | 1.012 |
| Square | Round | Neutral | 0.556 | 1.388 |
| Square | Square | Smiling | 1.116 | 1.072 |
| Square | Square | Neutral | 0.480 | 1.415 |
The ANOVA revealed a significant main effect of face shape, F(1, 21) = 5.562, p = 0.019, ηp² = 0.003. Participants gave higher liking scores to square-faced humanoid robots (M = 0.840, SE = 0.198) than to round-faced humanoid robots (M = 0.397, SE = 0.216). There was also a significant main effect of mouth expression, F(1, 21) = 19.640, p < 0.001, ηp² = 0.029. Smiling-mouth humanoid robots were rated as more likable (M = 0.944, SE = 0.179) than those with neutral linear mouths (M = 0.293, SE = 0.225). The main effect of eye shape was not significant, F(1, 21) = 0.645, p = 0.422, ηp² = 0.044. None of the two-way or three-way interactions reached significance (all p > 0.05). These subjective results suggest that face shape and mouth expression are of primary importance in the affective evaluation of humanoid robot faces.
Eye-Tracking Results
Fixation Time Distribution
We first examined how participants distributed their gaze across the AOIs. The average fixation duration in the eye region (236.26 ± 179.65 ms) was significantly longer than in the mouth region (89.69 ± 160.31 ms), t(21) = 6.41, p < 0.001. The face region exhibited a longer mean fixation time (261.15 ± 183.22 ms) than the mouth region, but the difference did not reach statistical significance (p = 0.068). This pattern indicates that, when viewing a humanoid robot face, participants primarily allocate their visual attention to the eyes and the general face area rather than the mouth.
Fixation Count
Table 2 lists the mean fixation counts for each of the eight humanoid robot face stimuli. The values were obtained by summing fixations across all AOIs (excluding background).
| Robot Stimulus | Mean Fixation Count | Standard Deviation |
|---|---|---|
| R1 | 21.78 | 5.39 |
| R2 | 22.00 | 3.43 |
| R3 | 22.13 | 3.19 |
| R4 | 25.79 | 3.85 |
| R5 | 22.83 | 3.82 |
| R6 | 26.92 | 4.23 |
| R7 | 22.29 | 2.81 |
| R8 | 22.88 | 3.87 |
A repeated-measures ANOVA on fixation count (Table 3) revealed a significant main effect of mouth expression, F(1, 21) = 5.395, p = 0.021, ηp² = 0.029. Participants produced more fixations when the humanoid robot had a smiling mouth (M = 25.12, SE = 1.44) than when it had a neutral linear mouth (M = 22.43, SE = 1.29). The main effect of eye shape was also significant, F(1, 21) = 8.504, p = 0.004, ηp² = 0.044. Round eyes attracted more fixations (M = 24.76, SE = 1.51) than square eyes (M = 22.79, SE = 1.44). There was no significant main effect of face shape, F(1, 21) = 0.477, p = 0.491, ηp² = 0.003.
| Source | df | F | p | ηp² |
|---|---|---|---|---|
| Face Shape | 1,21 | 0.477 | 0.491 | 0.003 |
| Mouth Expression | 1,21 | 5.395 | 0.021 | 0.029 |
| Eye Shape | 1,21 | 8.504 | 0.004 | 0.044 |
| Face Shape × Mouth Expression | 1,21 | 8.943 | 0.003 | 0.047 |
| Face Shape × Eye Shape | 1,21 | 0.833 | 0.362 | 0.005 |
| Mouth Expression × Eye Shape | 1,21 | 0.294 | 0.589 | 0.002 |
| Face Shape × Mouth Expression × Eye Shape | 1,21 | 0.179 | 0.673 | 0.001 |
The interaction between face shape and mouth expression was significant, F(1, 21) = 8.943, p = 0.003, ηp² = 0.047. Simple effects analysis showed that, when the humanoid robot had a square face, participants with a smiling mouth produced significantly more fixations than those with a neutral mouth (p = 0.008). Similarly, for round-eyed humanoid robots, the smiling mouth condition yielded more fixations than the neutral mouth condition (p < 0.001). These findings suggest that a smiling mouth, particularly in combination with a square face or round eyes, enhances visual attention allocation to the humanoid robot face.
EEG Results
Grand-average ERP waveforms for the eight humanoid robot face conditions are shown in Figure 2. We focus on the N1, P2, and P3 components in the centro-parietal and parieto-occipital regions. Figure 2 presents the time-locked waveforms (gray shading indicates the latency windows for each component). The mean amplitudes for each component across the eight stimuli are listed in Table 4.
| Robot Stimulus | N1 Amplitude (μV) | P2 Amplitude (μV) | P3 Amplitude (μV) |
|---|---|---|---|
| R1 | 2.88 | 6.42 | 4.96 |
| R2 | 2.37 | 7.16 | 4.94 |
| R3 | 3.73 | 4.40 | 5.00 |
| R4 | 5.18 | 4.28 | 4.01 |
| R5 | 1.73 | 7.68 | 4.83 |
| R6 | 1.68 | 7.17 | 5.38 |
| R7 | 3.98 | 4.68 | 5.02 |
| R8 | 5.03 | 6.11 | 6.19 |
For the N1 component, repeated-measures ANOVA (Table 5) revealed a highly significant main effect of face shape, F(1, 21) = 40.342, p < 0.001, ηp² = 0.632. Square-faced humanoid robots elicited a larger N1 amplitude (M = 4.30 μV, SE = 0.19) than round-faced humanoid robots (M = 2.49 μV, SE = 0.15). This indicates that square faces capture more attentional resources at the earliest perceptual stage (80–120 ms). The main effect of eye shape and mouth expression were not significant for N1 (p > 0.05). There was, however, a significant interaction between face shape and eye shape, F(1, 21) = 4.418, p = 0.042, ηp² = 0.099. Simple effects indicated that the face shape difference (square > round) was particularly pronounced when the eyes were round (p < 0.001).
| Component | Source | F | p | ηp² |
|---|---|---|---|---|
| N1 | Face Shape | 40.340 | <0.001 | 0.632 |
| N1 | Mouth Expression | 1.449 | 0.236 | 0.502 |
| N1 | Eye Shape | 1.780 | 0.190 | 0.035 |
| N1 | Face Shape × Eye Shape | 4.418 | 0.042 | 0.099 |
| P2 | Face Shape | 5.019 | 0.031 | 0.111 |
| P2 | Mouth Expression | 0.717 | 0.402 | 0.018 |
| P2 | Eye Shape | 0.147 | 0.703 | 0.004 |
| P3 | Mouth Expression | 4.906 | 0.033 | 0.001 |
| P3 | Mouth Expression × Eye Shape | 4.231 | 0.046 | 0.004 |
For the P2 component, the main effect of face shape was significant, F(1, 21) = 5.019, p = 0.031, ηp² = 0.111. Square-faced humanoid robots evoked a larger P2 amplitude (M = 6.86 μV, SE = 0.33) than round-faced humanoid robots (M = 5.02 μV, SE = 0.30). No other main effects or interactions were significant for P2 (all p > 0.05). The enhanced P2 contributes to the idea that square faces continue to dominate attention allocation during the later perceptual stage (150–250 ms).
For the P3 component, the main effect of mouth expression reached significance, F(1, 21) = 4.906, p = 0.033, ηp² = 0.001. Humanoid robots with a smiling mouth elicited a larger P3 amplitude (M = 5.24 μV, SE = 0.31) than those with a neutral linear mouth (M = 4.71 μV, SE = 0.27). Moreover, the interaction between mouth expression and eye shape was significant, F(1, 21) = 4.231, p = 0.046, ηp² = 0.004. Simple effects tests showed that, for smiling-mouth humanoid robots, round eyes produced a stronger P3 than square eyes (p = 0.033). In contrast, for neutral-mouth humanoid robots, eye shape did not significantly influence P3 (p > 0.05). This interaction indicates that a smiling mouth combined with round eyes intensifies the neural processing related to stimulus evaluation and attentional engagement at the P3 latency (300–450 ms).
Discussion
The present study provides converging evidence from subjective, oculomotor, and electrophysiological measures that specific facial features of humanoid robots systematically affect users’ affective perception. Our results confirm that the medium facial width-to-height ratio is more preferred than high or low ratios, echoing prior work suggesting that moderate proportions create a balance between dominance and approachability. This preference likely arises from the visual congruity between width-to-height ratio and personality attributions; a medium ratio avoids the overly serious impression conveyed by high ratios while preserving the friendly character associated with low ratios.
Face shape emerged as a robust determinant of both subjective liking and early neural responses. Participants consistently rated square-faced humanoid robots as more likable than round-faced ones. The ERP data showed that square faces evoked larger N1 and P2 amplitudes. The N1 component is often linked to early visual attention allocation to salient stimuli. The enhanced N1 for square faces suggests that their sharper geometric contours capture more initial attentional resources compared to softer round contours. Similarly, the P2 component reflects higher-order perceptual processing of task-relevant features. The larger P2 for square faces indicates that this feature continues to engage cognitive resources during the later stage of pattern recognition. These findings align with the notion that angular shapes in a humanoid robot face might be perceived as more futuristic or “robotic,” thereby gaining more focused visual attention from users.
Mouth expression had a consistent effect across subjective, eye-tracking, and EEG metrics. Smiling-mouth humanoid robots were not only rated as more likable but also received more fixations and elicited a larger P3 amplitude. The P3 component is widely believed to reflect the brain’s allocation of attentional resources for evaluating motivationally significant stimuli. A smiling mouth acts as a positive social cue, signaling friendliness and approachability, which likely increases the reward value of the humanoid robot face. Consequently, users devote more cognitive resources to processing the smiling mouth, as manifested in the elevated P3. The fact that the smiling mouth also increased fixation counts in the eye-tracking data reinforces the idea that this feature drives overt visual search and engagement.
The combined effect of smiling mouth and round eyes on P3 is particularly interesting. Round eyes alone did not produce a significant main effect on subjective liking or P3, but their interaction with mouth expression revealed that the enhancement of P3 was specific to the smiling-mouth context. This suggests that the combination of round eyes and a smiling mouth creates a more coherent and socially expressive humanoid robot face. The rounder eye shape may increase perceived anthropomorphism, while the smiling mouth provides a clear emotional signal; together they amplify the neural response associated with positive affect. Manufacturers of humanoid robots should therefore consider the holistic configuration of facial features rather than individual elements in isolation.
The eye-tracking results further revealed that participants spent more time looking at the eyes and face than at the mouth of humanoid robots. This is consistent with the natural human tendency to prioritize the eyes during facial processing. However, the significant effects of mouth expression and eye shape on fixation counts indicate that the mouth is not neglected; rather, a smiling mouth may draw additional fixations even if overall dwell time is lower. The interaction showing that more fixations occur when a square face is paired with a smiling mouth suggests that congruent design elements (e.g., an angular face with a clear positive emotional expression) synergistically boost visual attention.
Our findings have several practical implications for humanoid robot design. First, designers should use a medium facial width-to-height ratio as a baseline to ensure broad user appeal. Second, square face shapes may be advantageous for capturing attention and conveying a likable character, perhaps because they appear more assertive yet approachable. Third, incorporating a smiling mouth is essential for inducing positive emotions, as it increases both subjective liking and neural engagement. Finally, round eyes are recommended especially when a smiling mouth is present, because the combination maximizes the user’s cognitive and attentional processing, thereby facilitating a more positive human-robot interaction experience.
Conclusion
In this study, we used a multimodal approach combining subjective ratings, eye-tracking, and EEG to investigate how facial features of humanoid robots influence users’ affective perception. The main findings are: (1) a medium facial width-to-height ratio is preferred by users; (2) square face shapes evoke larger N1 and P2 amplitudes, indicating enhanced early attentional allocation, and are rated as more likable; (3) smiling mouths lead to higher fixation counts and stronger P3 amplitudes, reflecting increased cognitive and attentional engagement; (4) round eyes increase fixation counts compared to square eyes; and (5) the combination of a smiling mouth with round eyes elicits the strongest P3 response. These results demonstrate that objective physiological measures can complement subjective ratings to reveal the neural mechanisms underlying users’ affective responses to humanoid robot faces. Our study offers theoretical and practical support for designing humanoid robot faces that are not only aesthetically pleasing but also emotionally engaging, ultimately promoting smoother human-robot interactions.
