The Auditory Uncanny Valley (Voice & AI Assistants)

The Auditory Uncanny Valley (Voice & AI Assistants)

Summary

The Auditory Uncanny Valley extends Masahiro Mori's original hypothesis into the domain of voice and speech. As a synthetic voice becomes increasingly human-like — in tone, prosody, pacing, and conversational behaviour — listener affinity rises until a point where near-perfect but imperfect voices trigger discomfort, eeriness, or outright revulsion. The effect is amplified by real-time interactivity, making modern AI voice assistants a prime candidate for the valley.


From Visual to Auditory

Mori's 1970 curve mapped human likeness against affinity for visual appearance. The same logic applies to sound: a voice is a rich signal carrying identity, emotion, and intent. When a synthetic voice approaches but fails to reach human parity, the mismatch between expectation and reality produces the same cognitive dissonance described in The Uncanny Valley (不気味の谷現象).

The key difference is temporal: where visual uncanniness is often static or movement-based, voice is inherently dynamic — it unfolds over time, making the violation of expectation a moment-by-moment experience.


The Auditory Curve

Voice Type Human-likeness Affinity
Simple text-to-speech (e.g., early GPS, screen readers) Low Moderate (functional)
Stylised synthetic voice (e.g., Alexa, Siri's default) Medium High
Near-human neural TTS with subtle glitches High Low (the valley)
Actual human voice Perfect High

Mechanisms at Work

1. Predictive Coding & Expectation Violation

When a voice sounds almost human, the brain activates a full set of expectations: natural breath patterns, appropriate emotional inflection, correct rhythm and pacing. A neural TTS voice that stumbles on an unusual word, inserts an unnatural pause, or delivers a flat emotional response in a high-stakes context triggers a strong prediction error — the auditory equivalent of a humanoid robot's jerky movement.

2. Categorisation Ambiguity

A clearly robotic voice is safely categorised as "machine." A clearly human voice is "person." A voice that hovers in between — too expressive to be a tool, too artificial to be a person — resists easy categorisation, producing the same conflicting neural signals as a visual android.

3. Amplification by Interactivity

Mori theorised that movement amplifies the uncanny effect. In voice AI, real-time interactivity plays the same role. A static recording of a synthetic voice may be tolerable, but a conversational AI that hesitates, interrupts, laughs at the wrong moment, or changes tone mid-sentence falls much deeper into the valley. The interactive context raises the stakes on naturalness.


Real-World Examples


Is the Auditory Valley Shallower?

Some research suggests the auditory uncanny valley may be less severe than the visual one. Humans are accustomed to imperfect voices — phone call static, accents, speech impediments, emotional fatigue — so the tolerance band for vocal imperfection is wider. However, the effect becomes more pronounced as the AI's conversational behaviour (timing, turn-taking, empathy) approaches human levels, because behavioural uncanniness compounds acoustic uncanniness.


Design Implications


References