Study on the Structure and Control of a Humanoid Robot Head

My research journey into the field of humanoid robotics has been deeply focused on a specific yet highly challenging subsystem: the head. The humanoid robot head is not merely a mechanical enclosure for sensors; it is the primary interface for social interaction, perception, and expression. In this work, I present a comprehensive study covering the mechanical configuration, auditory localization, and control system of a humanoid robot head. The entire system was designed, manufactured, and tested within the framework of a university research project aiming at advancing humanoid robotics technologies.

The study of humanoid robots represents the pinnacle of modern robotics research, integrating mechanics, electronics, control theory, artificial intelligence, and biomimetics. Unlike industrial robots that operate in structured environments, humanoid robots are designed to interact with humans in unstructured, dynamic settings. The head, in particular, plays a critical role in this interaction, as it houses the primary sensory organs—vision and hearing—and provides the physical means for non-verbal communication such as nodding, turning, and expressing emotions. Therefore, the design of a humanoid robot head requires a careful balance between mechanical simplicity, functional versatility, and control precision.

1. Overall System Architecture

Before delving into the details of the head structure, I established a clear system architecture for the entire humanoid robot head platform. Based on an analysis of current technological capabilities and the specific requirements of our project, the system was divided into six major subsystems:

  1. Mechanical structure and actuation – the physical platform and motors
  2. Auditory localization – sound source direction detection
  3. Control system – motion control using servo motors and a multi-axis control card
  4. Vision processing – image acquisition and processing
  5. Olfactory processing – gas sensing capability
  6. Central processing unit – the main computer and coordination software

The work presented in this thesis focuses on the first three subsystems. The mechanical structure and actuation part acts as the “platform” for all other sensors and actuators. It is composed of lightweight aluminum alloy components with a compact “universal joint” (Cardan joint) configuration that allows the required degrees of freedom (DOF). The actuation is achieved through direct-drive DC servo motors, which eliminate the need for gearboxes and thus reduce backlash, friction, and nonlinearity, leading to better control performance.

For the auditory subsystem, two electret condenser microphones capture sound signals, which are then digitized by computer sound cards. The central processor performs cross-correlation analysis on the two channels to determine the time delay and hence the direction of the sound source. The control system uses a PMAC multi-axis control card (Delta Tau PMAC2A-PC/104) to generate accurate position and velocity commands for the servo motors. The control algorithm is based on PID with feedforward compensation, tuned through extensive experiments.

2. Mechanical Design of the Humanoid Robot Head

2.1 Determination of the Degrees of Freedom

Mimicking the full range of human head motion is a formidable task. The human head has many degrees of freedom, including the neck (three primary axes: yaw, pitch, and roll), the jaw (one DOF for opening/closing), and numerous facial muscles for expressions. For our project, we aimed to capture the essential motions required for natural human-robot interaction without excessive complexity. After careful consideration, I selected a configuration with four degrees of freedom:

  • Three DOFs in the neck: rotation (yaw), pitch (nodding), and roll (lateral tilt)
  • One DOF in the jaw: opening/closing for speech synchronization and expression

This configuration allows the robot head to perform all major head movements in a human-like manner. The jaw DOF is separated from the neck to provide a more natural two-stage nodding mechanism, which mirrors human anatomy more closely.

2.2 Structural Alternatives and Selection

Several structural schemes were considered to realize the three neck DOFs. These schemes differ in the arrangement of the rotational axes and the placement of motors. I systematically compared five different configurations, as summarized in Table 2.1 below.

Table 2.1 Comparison of Structural Schemes for the Neck
Scheme Description Advantages Disadvantages
1 Pitch axis and yaw axis not in the same plane; single cantilever support Simple construction Large bending moment on the support; requires high-torque motor
2 Pitch and yaw axes intersect (two-axis gimbal) Compact; no offset between axes One motor must support the entire structure; cable management issues
3 Motors placed in a non-orthogonal arrangement Flexible kinematics Cables may twist; complex wiring
4 All three rotational axes in a single plane Straightforward control logic Excessive overhang; large torque requirement for the first motor
5 Three non-collinear axes (two variants A and B) Balance between compactness and load distribution Complex manufacturing; wiring challenges

After analyzing these alternatives, I devised a structure based on a modified universal joint (Cardan joint) that provides two perpendicular rotational axes (pitch and roll) combined with a third motor for yaw rotation. This design achieves the three neck DOFs with a relatively simple and compact arrangement. The motor for yaw rotation is placed at the base of the neck, while the pitch and roll motors are integrated into the universal joint assembly. This arrangement minimizes the moving mass and reduces the required motor torque.

The final structural principle diagram is shown in Figure 2.1 (not reproduced here for brevity), which illustrates the four DOFs: three in the neck and one in the jaw. The jaw motion is achieved by a separate motor located inside the head shell, which drives a linkage mechanism to open and close the lower jaw.

Figure 2.1 illustrates the concept of quality inspection in the manufacturing of humanoid robot heads, emphasizing the need for precise assembly and testing.

2.3 Detailed Design and Specifications

The mechanical design was guided by principles of lightweight construction, structural symmetry, and low rotational inertia. The head shell was made of aluminum alloy, and the overall dimensions were set to approximate human head proportions: width 140 mm, depth 180 mm, and height 380 mm. The range of motion for each DOF is listed in Table 2.2.

Table 2.2 Motion Ranges of the Humanoid Robot Head
Motion Range
Primary pitch (nodding) -45° to +45°
Secondary pitch (jaw motion) -45° to +45°
Roll (lateral tilt) -45° to +45°
Yaw (head rotation) -90° to +90°

The load on the head includes two CCD cameras (approximately 0.43 kg each), the olfactory sensors, supporting structure, and the outer shell, totaling about 2 kg. The torque required to rotate this mass with an approximate lever arm of 0.1 m is about 2 N·m. This torque can be delivered by small, high-performance DC servo motors without the need for additional gearboxes. Therefore, I opted for a direct-drive configuration.

After completing the detailed mechanical drawings, the components were manufactured and assembled. The final assembled humanoid robot head is shown in Figure 2.2 (omitted here). The aluminum alloy structure proved to be sufficiently rigid while keeping the total weight low. Compared to other humanoid robot heads reported in the literature (e.g., the Kismet head from MIT or the WE-4 from Waseda University), my design has fewer degrees of freedom but is tailored specifically for neck and jaw movement rather than facial expressions. This intentional simplification allowed for a more robust and cost-effective solution that still meets the project’s objectives.

3. Sound Localization System

3.1 Principles of Auditory Localization

Human beings localize sound sources using three main cues: interaural time difference (ITD), interaural level difference (ILD), and spectral cues (or timbre difference). The head acts as an acoustic obstacle, causing frequency-dependent diffraction. For a sound source located at an angle θ relative to the perpendicular bisector of the line connecting the two ears, the ITD can be approximated by:

$$ \Delta t \approx \frac{2a \sin \theta}{c} $$

where a is the radius of the head (approximated as a sphere), and c is the speed of sound in air (approximately 340 m/s). If we define the effective distance between the two ears as h = 2a, the angle is:

$$ \theta \approx \arcsin\left(\frac{c \cdot \Delta t}{2a}\right) $$

The phase difference Δφ between the two ears is related to the ITD by:

$$ \Delta \phi = \omega \cdot \Delta t = \frac{2\pi f \cdot 2a \sin \theta}{c} $$

where f is the frequency. For low frequencies, the wavelength is large compared to the head size, so the ILD is small; the ITD plays the dominant role. For high frequencies, the head casts an acoustic shadow, and the ILD becomes significant. However, for our implementation, I relied primarily on ITD, as it is robust in the low-frequency range and can be accurately estimated using cross-correlation techniques.

3.2 Real-Time Audio Acquisition

For real-time sound detection, I chose to use two electret condenser microphones connected to two separate sound cards installed in the host computer. One sound card was the motherboard’s integrated sound card, and the other was a separate PCI sound card. To achieve low-latency capture, I employed Microsoft’s DirectX technology, specifically the DirectSound component. DirectSound provides a hardware abstraction layer (HAL) that allows direct access to audio hardware without requiring low-level programming. It also supports capture buffers, enabling continuous streaming of audio data.

The audio capture process involves:

  1. Enumerating audio capture devices using DirectSoundCaptureEnumerate.
  2. Creating capture buffer objects using CreateCaptureBuffer.
  3. Setting up circular (streaming) buffers and using notification events to read data.

Each capture buffer is circular, with two regions (upper and lower halves). A pointer indicates the current capture position, and event notifications trigger when the read pointer crosses certain boundaries. The application reads data from the buffer without needing to copy it to a separate processing buffer if the processing is fast enough.

The audio format was set to 22,050 Hz sampling rate, 8-bit resolution, and mono channel. The frame length for each processing window was chosen as 16,384 samples (2^14), which corresponds to about 0.743 seconds of audio. The capture buffer was set to two frames (32,768 bytes) to facilitate overlap and continuous processing.

3.3 Signal Processing and Direction Estimation

The two microphone signals, x(n) and y(n), are cross-correlated using the following formula:

$$ r_{xy}(m) = \sum_{n=0}^{N-1} x(n+m) \cdot y(n) $$

where N is the number of samples in the analysis window, and m is the lag (shift) in samples. By finding the value of m that maximizes r_xy(m), I obtain the interaural time delay in units of samples. The actual ITD is:

$$ \Delta t = \frac{m_{max}}{f_s} $$

where f_s is the sampling frequency (22,050 Hz). Substituting into the earlier angle formula gives the direction of the sound source.

The program flow for the sound localization system is summarized in Figure 3.1 (not shown). Initially, the program enumerates and initializes the sound capture devices, then launches separate threads for each channel to ensure simultaneous capture. The main thread handles the cross-correlation calculation and displays the resulting sound direction on a visual interface. Through extensive testing, I found that using the default 22.05 kHz sampling rate provided a good balance between accuracy and computational load. Higher rates increased precision but significantly impacted real-time performance.

Experimental results (Figure 3.2 in the original) demonstrated that the system could successfully estimate the direction of a sound source within a reasonable accuracy. The use of a standard PC sound card with DirectSound proved to be an economical and effective solution for real-time auditory localization on a humanoid robot head.

4. Control System Design

4.1 Hardware Components

The control system is the core of the humanoid robot head, ensuring precise and smooth motion. I selected the PMAC2A-PC/104 motion control card from Delta Tau as the main controller. This card is based on a Motorola 56300 DSP running at 40 MHz, with 128 KB SRAM and 512 KB flash memory. It can control up to 8 axes (with an expansion board) and supports various motor types including stepper, AC servo, and DC servo motors. Each axis provides a ±10 V analog command output, an encoder feedback input, and several digital I/O signals.

For actuation, I chose DC servo motors from Maxon Motor. Specifically, the RE 025-055-35EBAZ01A motor with a planetary gearhead (gear ratio 128:1) and an incremental encoder (50 counts per revolution, HEDS-5540 encoder). The motor has a nominal power of 20 W, a maximum continuous torque of 1.3 N·m, and a rated speed of 9550 rpm. The gearhead reduces the speed to an appropriate range while increasing the output torque to meet the ~2 N·m requirement.

The servo amplifiers are Copley Controls Accelnet MicroPanel series, model ACJ-055-09, with a maximum current of 9 A and a DC supply voltage of 5 V. These amplifiers accept analog command signals from the PMAC card and provide current feedback to the motor. To condition the encoder signals, a signal adapter box was built using a MC3487N differential line driver, which converts the single-ended encoder outputs into differential signals suitable for the amplifier’s inputs.

4.2 Tuning and Calibration

The first step was to connect the motor and amplifier, then use the CMEZ software provided by Copley to set the current loop and velocity loop parameters. The CMEZ software communicates via RS-232 and offers both automatic and manual tuning modes. I first ran the automatic tuning for the current loop (as shown in Figure 4.1 in the original), then used the oscilloscope function to fine-tune the proportional and integral gains to minimize the error between the command and actual current. The final gains were stored in the amplifier’s flash memory.

Subsequently, the PMAC card was connected to the amplifier and the host computer via a serial port. The PMAC configuration variables were set as follows:

Table 4.1 PMAC Configuration Variables
Variable Value Description
I900 1001 PWM frequency for channels 1-4 (kHz)
I901 2 Phase clock frequency divider
I902 3 Servo clock frequency divider
I906 1001 PWM frequency for channels 5-8
I916 0 Phase clock selection
I969 1024 Maximum DAC output
I100 1710933 Interpolated time between servo cycles

The effective servo clock frequency was set to 4 kHz, which is typical for high-performance motion control. The DAC output range was set to ±10 V via I969.

Before enabling closed-loop control, I verified the polarity of the feedback. A simple open-loop command (e.g., O40) should cause the position counter to increase; a negative command should decrease it. If the polarity was reversed, I adjusted I900 (or the wiring) to correct it. After confirming correct polarity, I set the proportional gain (I130) to 200 initially, then adjusted the derivative gain (I131) to eliminate oscillations. The integral gain (I132) was set as needed for zero steady-state error.

4.3 Visual Control Software in VC++

For the high-level control, I developed a Windows-based application using Visual C++ 6.0 with the MFC (Microsoft Foundation Classes) library. The program communicates with the PMAC card through the PTalkDT ActiveX control provided by Delta Tau. This control encapsulates the low-level communication and allows the sending of PMAC commands and receiving responses. The key method is GetResponse(response, command), which sends a command string to PMAC and waits for the response.

To control a single motor to rotate to a specific angle, the program calculates the required number of encoder pulses. The motor revolution count was determined as 256,000 pulses per revolution (50 lines × 4 for quadrature × 128 gear ratio). The user enters the desired angle and a scale factor; the program computes the command:

PulseCount = (Angle × ScaleFactor) / 360

Then the command string "J: " + PulseCount is sent via GetResponse. A “home” command is first issued to initialize the axis. The graphical user interface (shown in Figure 4.2) displays the current position, command status, and allows real-time adjustment of parameters. For four-axis control (two cameras, neck yaw/pitch/roll, and jaw), a similar interface was developed with independent control for each axis.

The program flow is as follows:

  1. Initialize the PTalkDT control and establish communication with PMAC.
  2. Read user input (angle, scale factor).
  3. Calculate pulse count and format command.
  4. Send “home” command to zero the axis.
  5. Send “J :” command with pulse count.
  6. Monitor position feedback and display status.

Extensive testing was conducted to coordinate the feedback path (encoder) and command path. The PID parameters were tuned to achieve a fast, stable response with minimal overshoot. The final system achieved precise angular positioning within the tolerance of the encoder resolution (0.0014 degrees per pulse after quadrature and gear ratio).

5. Summary and Future Work

In this thesis, I have presented a comprehensive study on the structure, auditory localization, and control system of a humanoid robot head. The main contributions are:

  • A novel four-DOF head mechanism (three in the neck, one in the jaw) based on a universal joint configuration. The design is lightweight, compact, and implements direct-drive servo motors for high accuracy.
  • A real-time sound localization system using two microphones and a standard PC sound card. Cross-correlation of the two signals provides the ITD, allowing the estimation of sound source direction. The system was implemented and validated experimentally.
  • A complete motion control system based on a PMAC multi-axis control card and DC servo motors. PID parameters were tuned and validated through extensive experiments. A user-friendly VC++ interface was developed for precise position control.

The research has laid a solid foundation for further development of the humanoid robot head. However, several challenges remain:

  • The sound localization accuracy can be improved by incorporating higher sampling rates, advanced signal processing (e.g., generalized cross-correlation with phase transform, GCC-PHAT), and considering reverberation and noise robustness.
  • The control system could benefit from feedforward compensation and adaptive control algorithms to handle varying loads and disturbances.
  • The integration of vision and auditory systems is essential for a truly interactive humanoid robot head. Future work should focus on sensor fusion and coordinated control.

Overall, the humanoid robot head system developed in this work demonstrates the feasibility of using relatively simple and cost-effective components to achieve sophisticated behaviors. The knowledge gained will guide future enhancements and the development of a complete humanoid robot.

In closing, I would like to emphasize that the field of humanoid robotics remains an exciting and challenging frontier. Each subsystem—structure, perception, and control—presents unique problems that require creative engineering solutions. My study on the humanoid robot head is a small step toward the ultimate goal of creating machines that can coexist and interact with humans naturally and meaningfully.

Scroll to Top