Affinity-Driven Design of Humanoid Robots via Stable Diffusion Model Training

In the rapidly evolving landscape of robotics, the humanoid robot has emerged as a central figure in both domestic and commercial service sectors. As a designer and researcher deeply engaged in human-robot interaction, I have observed that the aesthetic and emotional qualities of a humanoid robot significantly influence user acceptance and trust. Many existing humanoid robots are engineered from a purely functional perspective, resulting in cold, rigid, and unapproachable appearances. This lack of affinity often triggers user resistance and reduces the willingness to engage with these machines. Therefore, my work focuses on systematically exploring the affinity-related design features of humanoid robots and leveraging artificial intelligence (AI) to generate more appealing design alternatives. In this paper, I present a comprehensive methodology that integrates kansei engineering with stable diffusion (SD) model training to achieve this goal.

The core of my approach lies in quantifying the abstract concept of “affinity” into concrete design elements. Through kansei engineering, I established a clear mapping between user emotional responses and the visual attributes of humanoid robots, including form, material, texture, and color. This mapping enables me to construct a precise scoring framework for evaluating and generating high-affinity designs. Furthermore, I developed a specialized training pipeline for the SD model, allowing it to produce a wide array of humanoid robot concepts that are not only diverse but also consistently aligned with the identified affinity criteria. This integration of perceptual engineering and generative AI represents a novel and effective path for humanoid robot design, significantly enhancing both design efficiency and the emotional quality of the final products.

Kansei Engineering for Humanoid Robot Affinity

To systematically explore the affinity characteristics of humanoid robot appearances, I adopted the kansei engineering methodology. This approach is particularly effective in translating human emotional responses into concrete engineering parameters. I began by collecting a comprehensive set of image word pairs related to affinity. After filtration and expert review, I selected three primary evaluation dimensions: Affinity (cold-alienating vs. warm-approachable), Tenderness (hard-cold vs. soft-gentle), and Liveliness (serious-rigid vs. playful-energetic). These dimensions collectively represent the perceived emotional quality of a humanoid robot.

I designed a questionnaire using abstracted visual representations of humanoid robot components to isolate the effect of individual design elements. The questionnaire covered five key categories: head shape, eye shape, body proportion, material type and surface finish, and color scheme. For each category, I generated multiple variants. For example, head shapes included square, circle, vertical rounded rectangle, horizontal rounded rectangle, and semicircle. Eye shapes included circle, square, vertical racetrack, and horizontal racetrack. Body proportions ranged from infantile to adult, with both slender and sturdy silhouettes. Materials included plastic, fabric, metal, and transparent materials, each with two surface treatments. Colors were carefully selected across warm, cool, and neutral palettes, with variations in saturation and brightness.

I collected 645 valid responses and performed extensive statistical analyses. The results revealed that the influence of design elements on perceived affinity, in descending order, was: head shape, facial expression, color, material, and body proportion. To understand the inter-relationships among the three affinity dimensions, I conducted correlation analyses. The results are summarized in Table 1 and Table 2 below.

Table 1: Kendall correlation coefficients among affinity dimensions for categorical design elements

Design element Affinity vs Tenderness Affinity vs Liveliness Tenderness vs Liveliness
Head shape 0.356** 0.303** 0.284**
Eye shape 0.318** 0.345** 0.368**
Body proportion 0.160** 0.113** 0.152**
Plastic 0.048** -0.006 -0.060**
Fabric -0.016 -0.006 -0.020
Metal 0.029 -0.010 0.028
Transparent -0.001 0.013 -0.032
White/Gray/Black 0.294** 0.319** 0.301**
High saturation warm 0.055** 0.006 0.038*
Low saturation warm -0.020 0.004 -0.003
High saturation cool 0.047** -0.020 -0.007
Low saturation cool 0.032 0.006 -0.019

Note: * indicates P < 0.05, ** indicates P < 0.01. The results show strong positive correlations among the three dimensions, confirming that they collectively describe the concept of affinity.

Table 2: Pearson correlation coefficients among affinity dimensions for continuous/quantitative design elements

Design element group Affinity vs Tenderness Affinity vs Liveliness Tenderness vs Liveliness
Four material types 0.501** 0.148** 0.143**
Warm colors high vs low saturation 0.144** -0.062* 0.007
Cool colors high vs low saturation 0.121** 0.016 0.035
White vs high saturation 0.280** 0.158** 0.180**
White vs low saturation 0.155** 0.128** 0.134**

These correlations confirm that the three evaluation dimensions are strongly consistent, meaning that a design which is perceived as affinitive is also typically perceived as tender and lively. Interestingly, for surface treatments of materials and for high-saturation warm colors, the relationship between affinity and liveliness was negative. This suggests that enhancing the liveliness of a humanoid robot may sometimes reduce its perceived affability, indicating a delicate balance in design decisions.

Next, I performed non-parametric tests (Kruskal-Wallis H and Mann-Whitney U) to evaluate the specific contributions of each design variant. Table 3 presents the median scores (with interquartile ranges) for head shapes, eye shapes, and body proportions. Scores above 3 indicate a positive perception.

Table 3: Kruskal-Wallis H test results for head shape, eye shape, and body proportion

Design element Category Affinity M(P25, P75) Tenderness M(P25, P75) Liveliness M(P25, P75)
Head shape A1 (square) 6 (5, 7) 6 (5, 7) 6 (5, 7)
A2 (circle) 6 (5, 6) 6 (5, 6) 6 (5, 6)
A3 (vertical rounded rect) 3 (2, 4) 3 (2, 4) 4 (3, 5)
A4 (horizontal rounded rect) 6 (5, 7) 6 (5, 7) 6 (5, 7)
A5 (semicircle) 3 (2, 4) 3 (2.5, 4.5) 4 (3, 5)
Eye shape B1 (circle) 4 (4, 5) 4 (3, 5) 4 (3, 5)
B2 (square) 4 (4, 5) 4 (4, 5) 4 (4, 5)
B3 (vertical rounded rect) 6 (5, 7) 6 (5, 7) 6 (5, 7)
B4 (horizontal rounded rect) 3 (2, 3) 3 (2, 3) 2 (2, 3)
Body proportion C1 (infantile slender) 6 (5, 6) 5 (5, 6) 5 (5, 6)
C2 (child sturdy) 5 (4, 6) 5 (4, 6) 5 (4, 6)
C3 (child slender) 6 (5, 7) 6 (5, 7) 6 (5, 6)
C4 (teen slender) 6 (4, 6) 5 (5, 6) 5 (4, 6)
C5 (teen sturdy) 5 (4, 6) 6 (5, 6) 5 (3, 5)
C6 (adult sturdy) 4 (3, 5) 4 (3, 5) 4 (3, 5)
C7 (adult very sturdy) 4 (3, 5) 3 (2, 4) 4 (3, 5)

From the results, I found that the horizontal rounded rectangle (A4) head shape was perceived as most affinitive, followed by the circle (A2). The vertical rounded rectangle eye shape (B3) scored highest, while the horizontal racetrack eye (B4) was deemed non-affinitive. For body proportions, the infantile slender (C1) and child slender (C3) figures received the highest affinity scores, while adult sturdy figures (C6, C7) were rated lowest. These findings indicate that a humanoid robot with a cute, rounded facial contour and a youthful, slender body is more likely to evoke positive emotional responses.

Table 4: Mann-Whitney U test results for material and surface finish

Design element Category Affinity M(P25, P75) Tenderness M(P25, P75) Liveliness M(P25, P75)
Plastic D1 (smooth glossy) 4 (3, 5) 4 (3, 5) 5 (4, 5)
D2 (matte) 5 (4, 5) 5 (4, 5) 4 (3, 5)
Fabric D3 (plush) 6 (5, 7) 6 (5, 7) 5 (4, 6)
D4 (woven) 6 (5, 7) 6 (4, 6) 5 (4, 6)
Metal D5 (polished) 2 (1, 3) 2 (1, 3) 4 (2, 5)
D6 (brushed) 2 (1, 4) 2 (1, 4) 4 (3, 6)
Transparent D7 (clear smooth) 4 (3, 5) 4 (2, 5) 5 (4, 6)
D8 (frosted) 4 (4, 5) 4 (3, 5) 4 (3, 5)

Material analysis clearly showed that fabric, especially plush fabric, yielded the highest affinity scores, while metal scored the lowest. Matte surfaces were more affinitive than glossy ones for both plastic and transparent materials. This suggests that to increase the affinity of a humanoid robot, designers should favor soft, textured materials and avoid cold, reflective surfaces.

For color analysis, I employed Spearman correlation to examine the effects of height, neutral color value, and hue position within warm and cool groups. Table 5 lists the correlation coefficients.

Table 5: Spearman correlation coefficients within design element groups

Design element change Affinity Tenderness Liveliness
Overall height (low to high) -0.340** -0.427** -0.340**
Slender body height (low to high) -0.011 -0.037 -0.089**
Sturdy body height (low to high) -0.379** -0.541** -0.355**
Neutral color (white to black) -0.613** -0.622** -0.633**
High saturation warm (right to left) 0.102** 0.068** 0.054**
Low saturation warm (right to left) 0.071** 0.095** 0.095**
High saturation cool (left to right) 0.220** 0.228** -0.072**
Low saturation cool (left to right) 0.067** 0.145** -0.078**

These results demonstrate that taller bodies, especially sturdy ones, strongly reduce affinity. White is the most affinitive neutral color, while black is the least. In the warm hue spectrum, colors positioned further toward the red-orange side (right to left in the HSB diagram) increase affinity. In the cool spectrum, colors toward the green-cyan side (left to right) are more affinitive. Additionally, I performed linear regression to compare the effects of saturation levels. The coefficients in Table 6 indicate the direction and magnitude of differences, with positive values meaning the latter group has a stronger positive effect on the given dimension.

Table 6: Linear regression coefficients between color groups

Comparison group Affinity Tenderness Liveliness
High vs low saturation (warm) 0.867 0.672 -0.208
High vs low saturation (cool) 0.855 0.668 -0.012
White vs high saturation cool -1.633 -1.647 -0.705
White vs high saturation warm -0.466 -0.273 0.421
White vs low saturation cool -0.778 -0.979 -0.717
White vs low saturation warm 0.401 0.399 0.213

The regression results reveal that low-saturation colors universally outperform high-saturation colors in terms of affinity. The most affinitive color was low-saturation warm, followed by white, then high-saturation warm, low-saturation cool, and finally high-saturation cool. Therefore, for humanoid robot design, I recommend using soft, pastel-like colors with high brightness and low saturation. Warm whites, cream, and light orange are excellent choices to create a welcoming and comfortable appearance.

Development of an Affinity Scoring Table for Humanoid Robot Design

Based on the comprehensive statistical analyses, I constructed a quantitative affinity scoring table for each design element. This table serves as the fundamental reference for both evaluating existing humanoid robot designs and guiding the generation of new concepts through AI. The scoring table is presented below. Each design feature is assigned an integer score from 0 (low affinity) to 4 (high affinity). For categories with multiple sub-features, the scores are aggregated according to their relative importance and statistical results. The maximum possible total score for a design is 16.

Table 7: Affinity scoring table for humanoid robot design features

Element Feature Score Element Feature Score
Head shape A1 square 3 Eye shape B1 circle 1
A2 circle 3 B2 square 1
A3 vertical rounded rect 0 B3 vertical rounded rect 3
A4 horizontal rounded rect 3 B4 horizontal rounded rect 0
A5 semicircle 0
Body proportion C1 infantile slender 3 Material type D1 plastic glossy 1
C2 child sturdy 2 D2 plastic matte 2
C3 child slender 4 D3 fabric plush 3
C4 teen slender 3 D4 fabric woven 3
C5 teen sturdy 2 D5 metal polished 0
C6 adult sturdy 1 D6 metal brushed 0
C7 adult very sturdy 0 D7 transparent glossy 1
D8 transparent frosted 1
Color E1 low saturation warm 4
E2 high saturation warm 2
E3 low saturation cool 2
E4 high saturation cool 1
E5 white 3
E6 gray 0
E7 black 0

This scoring table encapsulates my empirical findings. For example, a humanoid robot with a horizontal rounded head (A4), vertical rounded eyes (B3), child slender body (C3), fabric plush material (D3), and low-saturation warm color (E1) would achieve the maximum score of 16. In contrast, a robot with a semicircular head, horizontal eyes, adult very sturdy body, polished metal, and black color would receive a score of 0, indicating a total lack of affinity.

Training the Stable Diffusion Model for Affinity Generation

With a robust quantitative foundation in place, I proceeded to train the stable diffusion model to generate high-affinity humanoid robot designs. The stable diffusion model is a powerful generative AI framework that can synthesize images based on text prompts. However, relying solely on prompt engineering often fails to produce designs that precisely match the nuanced affinity criteria identified in my study. Therefore, I developed a structured training protocol involving two complementary techniques: Dreambooth and Low-Rank Adaptation (LoRA). Dreambooth fine-tunes all layers of the neural network, enabling the model to learn specific object categories and styles. LoRA, on the other hand, inserts lightweight adaptation layers, offering a more efficient and compact way to influence generation. By combining both, I aimed to achieve both high fidelity and versatility in generating humanoid robot appearances.

The training process consisted of four main stages: sample preparation, sample annotation, iterative training, and model evaluation and deployment. Each stage is described in detail below.

2.1 Training Sample Preparation

The quality of the training samples directly determines the output quality of the model. I curated a collection of humanoid robot images that satisfied the affinity criteria derived from the scoring table. The samples were selected to be stylistically consistent yet rich in detail variety, ensuring that the model would learn the underlying affinity features rather than overfit to specific instances. To further enhance the samples, I utilized several stable diffusion modules such as ControlNet, image-to-image inpainting, and local repainting, alongside manual adjustments in Photoshop. This iterative refinement process allowed me to progressively align the samples with the desired affinity traits. An example of such refinement involved modifying a robot’s head shape from a less affinitive vertical rounded rectangle to a highly affinitive horizontal rounded rectangle, adjusting eye shape from horizontal to vertical rounded rectangle, and changing the body proportion from a sturdy adult to a slender child.

Each refinement step was guided by the affinity scoring table. For instance, I would score each candidate sample according to Table 7 and prioritize those with higher scores. This data-driven approach ensured that the training dataset was not only visually appealing but also statistically aligned with user preferences.

2.2 Training Sample Annotation

Proper annotation is crucial for text-to-image diffusion models. I first used an automatic annotation tool to generate base captions for each image. Then, I manually refined these annotations to precisely describe the affinity-relevant design features. The manual annotation covered the following aspects: overall silhouette, detailed facial features, stylistic descriptors, material and surface texture, and color palette. For example, a well-designed affinitive humanoid robot might be annotated as:

“affinity robot, minimalism, organic form, science fiction, humanoid robot, oval head, vertical eyes, yellow eyes, no mouth, white body, plastic material, shiny joints, teenage figure, full body, standing, arms at sides.”

This annotation describes not only the visual components but also the stylistic and emotional tone, providing the model with rich semantic guidance for generating similar designs.

2.3 Iterative Training and Refinement

I conducted multiple training rounds to iteratively improve the model’s output quality. In each round, I generated a set of designs and evaluated them against the affinity scoring table. If the generated humanoid robot designs deviated from the desired affinity characteristics, I adjusted the training samples accordingly and retrained the model. Table 8 shows the progression across three representative rounds.

Table 8: Affinity scores of representative training samples across rounds

Round Head shape Eye shape Body proportion Material type Overall color Total score
1 A4 (horizontal rounded rect) B2 (square) C3 (child slender) D1 (plastic glossy) E1 (low sat warm) 12
2 A4 B3 (vertical rounded rect) C1 (infantile slender) D2 (plastic matte) E5 (white) 14
3 A4 B3 C3 (child slender) D2 (plastic matte) E1 (low sat warm) 16

In the first round, the generated designs had appropriate head and eye shapes but appeared too cartoonish and lacked realism. To address this, I introduced more photorealistic samples in the second round. However, this caused the body proportions to become overly short and infantile. In the third round, I corrected the body proportions to the highly scored child slender type and incorporated more diverse design details. The resulting outputs from this round closely matched the desired affinity characteristics.

The training parameters were fine-tuned through experimentation. The optimal settings are listed in Table 9.

Table 9: Training hyperparameters for the stable diffusion model

Parameter Value Parameter Value
Learning Rate 1×10⁻⁴ Optimizer 8bit-Adam
Iteration 10 Scheduler Cosine
Batch Size 5 DIM (LoRA dimension) 128
Epoch 10 Alpha (LoRA alpha) 64

These parameters provided a good balance between training speed and output quality. The learning rate was kept relatively low to prevent catastrophic forgetting, while the cosine scheduler ensured smooth convergence. The LoRA dimension of 128 and alpha of 64 allowed for sufficient expressiveness without excessively increasing model size.

2.4 Model Evaluation and Selection

After each training epoch, the model checkpoint was saved. The loss value (Loss) served as a preliminary quantitative indicator of training quality. An ideal loss curve typically decreases initially, then begins to rise after a certain point due to overfitting. In my experiments, I observed that the loss reached approximately 0.08 between the 7th and 9th epochs, which I considered optimal. I used this criterion to select promising checkpoints for further evaluation.

To validate the practical effect of different checkpoints and model weights, I generated an XY cross-grid plot. The horizontal axis represented different model weights (0.2 to 1.0), and the vertical axis represented different epochs (3, 6, 9). Each cell contained a generated humanoid robot image. I then scored each image using the affinity scoring table. Table 10 shows the aggregate scores for each combination.

Table 10: Affinity scores for XY cross-grid (model weight vs epoch)

Epoch Loss value Weight 0.2 Weight 0.4 Weight 0.6 Weight 0.8 Weight 1.0
3 0.089 9 13 13 13 13
6 0.075 9 9 14 14 13
9 0.068 9 11 12 12 12

The best performance was achieved at epoch 6 with a model weight of 0.6 to 0.8, producing humanoid robot concepts with an affinity score of 14. This configuration balanced the influence of the base model and the fine-tuned features, yielding designs that were both realistic and emotionally appealing.

Finally, I combined the optimal Dreambooth model with multiple LoRA style models to generate a diverse style matrix. Each style model contributed a unique aesthetic (e.g., cartoonish, minimalist, futuristic, etc.) while preserving the core affinity features. I generated five design variants for each of four styles. Table 11 lists the affinity scores for these variants.

Table 11: Affinity scores of the style matrix

Style ID Variant 1 Variant 2 Variant 3 Variant 4 Variant 5
1 15 12 10 9 8
2 15 14 14 14 14
3 15 14 14 12 12
4 14 12 12 9 9

The first variant in each style consistently achieved the highest affinity score, often reaching the maximum of 15. This indicates that the model can reliably generate high-quality humanoid robot designs with strong emotional appeal. The style matrix allows designers to explore a wide range of visual directions while maintaining the core affinity attributes.

Mathematical Formulation of the Design Evaluation

To provide a more rigorous framework, I formulated the affinity score of a humanoid robot design as a weighted sum of its component scores. Let \( S \) denote the total affinity score, and let \(s_{\text{head}}\), \(s_{\text{eyes}}\), \(s_{\text{body}}\), \(s_{\text{material}}\), and \(s_{\text{color}}\) be the scores from Table 7. The total score is computed as:

$$ S = s_{\text{head}} + s_{\text{eyes}} + s_{\text{body}} + s_{\text{material}} + s_{\text{color}} $$

where each sub-score ranges from 0 to a maximum value (4 for body and color, 3 for head and material, 3 for eyes). The maximum total score is 16. This linear formulation is simple and effective for ranking design alternatives. However, based on the correlation analysis, I know that some dimensions interact. For instance, the negative correlation between liveliness and affinity for certain materials suggests that an overly lively design might reduce affinity. In future work, I plan to explore non-linear models, such as:

$$ S = \sum_{i=1}^{n} w_i s_i + \sum_{i<j} $$="" \(="" \)="" accurate="" analysis.="" and="" are="" capture="" derived="" effects="" even="" from="" interaction="" more="" of="" p="" perception.

During model training, the objective is to minimize the diffusion loss function. In the latent diffusion model, the training objective is:

$$ L = \mathbb{E}_{x, \epsilon, t} \left[ \left\| \epsilon – \epsilon_\theta (z_t, t, c) \right\|_2^2 \right] $$

where \( x \) is the input image, \( \epsilon \) is the noise added at timestep \( t \), \( z_t \) is the latent representation of the noisy image, \( c \) is the conditioning text embedding, and \( \epsilon_\theta \) is the denoising network. By training the model on a dataset of high-affinity humanoid robot images with corresponding annotations, the network learns to generate images that align with the semantic concept of “affinity” in the context of humanoid robots.

For LoRA, the update to the weight matrix \( W \) is constrained to be low-rank:

$$ W’ = W + \Delta W = W + B A $$

where \( B \in \mathbb{R}^{d \times r} \), \( A \in \mathbb{R}^{r \times k} \), and \( r \ll \min(d,k) \). This reduces the number of trainable parameters significantly while still allowing effective fine-tuning. In my implementation, I set \( r = 128 \) (DIM) and \( \alpha = 64 \) (Alpha), which balances capacity and efficiency.

Conclusion and Future Directions

In this paper, I presented a comprehensive methodology for exploring and generating affinity-driven humanoid robot appearances. By integrating kansei engineering with stable diffusion model training, I achieved several key contributions. First, I quantitatively revealed the design characteristics that enhance the affinity of a humanoid robot. These include horizontal rounded head shapes, vertical rounded eye shapes, slender youthful body proportions, soft textured materials such as fabric and matte plastic, and low-saturation warm colors. Second, I established an affinity scoring table that provides a precise and actionable guideline for designers and AI models alike. Third, I developed a robust training pipeline that successfully fine-tunes a stable diffusion model to generate a wide variety of high-affinity humanoid robot designs. The combination of Dreambooth and LoRA enabled both high-quality generation and stylistic diversity. The resulting style matrix demonstrates the model’s ability to produce multiple creative directions while consistently maintaining high affinity scores.

The implications of this work extend beyond aesthetics. A humanoid robot with a high-affinity appearance can significantly improve user acceptance and trust, which is critical for the widespread deployment of humanoid robots in homes, hospitals, and retail environments. My approach provides a data-driven and AI-accelerated path for achieving this goal, reducing the time and effort required for iterative design.

However, I acknowledge certain limitations. The questionnaire respondents were predominantly from a specific cultural background, which may affect the generalizability of the findings. Affinity perceptions can vary across cultures, and future studies should involve a more diverse participant pool. Additionally, my current model focuses on static appearance only. The affinity of a humanoid robot is also influenced by dynamic factors such as facial expressions, gestures, and voice. Integrating these interactive elements into the generative model would create an even more holistic design tool.

In future work, I plan to expand the research in three main directions. First, I will conduct cross-cultural studies to refine the affinity scoring table for different user groups. Second, I aim to incorporate style diversification techniques that allow the model to explore more unconventional yet still affinitive designs. Third, I will investigate the synergy between appearance and interaction behavior, using multimodal AI models to generate humanoid robots that are not only visually appealing but also behaviorally engaging. These efforts will further solidify the role of AI-assisted design in creating humanoid robots that are truly accepted and loved by users.

Overall, this study demonstrates that the combination of kansei engineering and stable diffusion model training offers a powerful and efficient methodology for humanoid robot design. It bridges the gap between user emotions and AI generation, enabling designers to create humanoid robots that are not just functional machines but also warm and approachable companions. As the field of humanoid robotics continues to advance, I believe that such data-driven, human-centered design approaches will become increasingly essential to ensuring that these machines are integrated seamlessly into our daily lives.

Scroll to Top