Local Non-Modal Voice Quality Dynamics in Japanese Anime Speech

Abstract

Japanese anime voices exhibit distinctive expressive characteristics shaped by the culture of professional voice actors (seiyus). Prior research has focused mainly on global phonetic properties, leaving local, dynamic voice quality changes underexplored. This study analyzes non-modal voice qualities—breathiness, harshness, pressed voice, and inhalations— in performances by professional seiyus and seiyu-training students across dubbing, repeating, and reading tasks. Professionals showed a high rate of non-modal segments, whereas students and reading-style speech produced fewer, confirming that local voice modulation is central to anime performance. Breathiness appeared mainly at phrase-final positions, while harsh and pressed qualities were more common phrase-initially and medially. Acoustic measures (HNR, F1F3syn, H1–A1) reliably distinguished the voice quality categories and were consistent across speaker groups. Overall, the findings highlight local voice quality modulation as a key expressive feature of anime voice acting.

Explanation about the samples

Speech samples are provided for the segmented non-modal voice quality categories. A margin of 100ms before the target segment is included to provide phonetic context. A pause of 200ms is included between the segments. The amplitudes of the segments are normalized to allow easier listening.

Speech samples dubbing task

aspiration

aspirated voice

breathy voice

harsh voice

harsh whispery

pressed voice

unvoiced inhalation

voiced inhalation