Abstract
Japanese anime voices exhibit distinctive expressive characteristics shaped by the culture of professional voice actors (seiyus). Prior research has focused mainly on global phonetic properties, leaving local, dynamic voice quality changes underexplored. This study analyzes non-modal voice qualities—breathiness, harshness, pressed voice, and inhalations— in performances by professional seiyus and seiyu-training students across dubbing, repeating, and reading tasks. Professionals showed a high rate of non-modal segments, whereas students and reading-style speech produced fewer, confirming that local voice modulation is central to anime performance. Breathiness appeared mainly at phrase-final positions, while harsh and pressed qualities were more common phrase-initially and medially. Acoustic measures (HNR, F1F3syn, H1–A1) reliably distinguished the voice quality categories and were consistent across speaker groups. Overall, the findings highlight local voice quality modulation as a key expressive feature of anime voice acting.
Explanation about the samples
Speech samples are provided for the segmented non-modal voice quality categories. A margin of 100ms before the target segment is included to provide phonetic context. A pause of 200ms is included between the segments. The amplitudes of the segments are normalized to allow easier listening.
| Speech samples | dubbing task |
|---|---|
|
aspiration |
|
|
aspirated voice |
|
|
breathy voice |
|
|
harsh voice |
|
|
harsh whispery |
|
|
pressed voice |
|
|
unvoiced inhalation |
|
|
voiced inhalation |