The voice of a virtual human is a powerful cue to its identity, personality, age, emotion and social characteristics. Psychological research on cross-modal face–voice perception suggests that people have expectations about which voices belong with which faces. People can make associations between unfamiliar faces and voices, while characteristics such as vocal pitch, apparent age, femininity/masculinity, prosody and expressiveness can influence the facial characteristics we associate with a voice. Lander et al. (2007) showed that dynamic aspects of speaking style contribute to face–voice identity matching, while other work has demonstrated systematic associations between vocal and facial age and femininity/masculinity.
In this project, the student will create a set of speaking virtual human characters and use modern generative speech tools to systematically manipulate properties of their voices, such as pitch, apparent age, speaking style, expressiveness or vocal identity. Virtual character appearance can similarly be manipulated along dimensions suggested by the psychology literature, allowing the same character and dialogue to be presented with voices that are either congruent or incongruent with its appearance.
A perceptual experiment will investigate which combinations of facial and vocal characteristics appear to “belong together”, and whether mismatches affect perceived personality, naturalness, trustworthiness, social presence or character identity. The project could additionally investigate whether these effects depend on visual realism—for example, whether observers are more sensitive to an incongruent voice for a highly realistic MetaHuman than for a stylised character.
The project combines speech synthesis, generative AI, virtual humans, Unreal Engine 5 and human perception, and would suit a student interested in speech/audio, AI and real-time character technology.