I bet you're talking about like English dubs. The original japanese voices are diverse. Aizen from Bleach, Hakshaku (Millennium Earl) from D. Grayman, Marshal D. Teach from One-piece, and All Might from my hero academia have deep regal voices.
The voice actors usually try to match the timing of the mouth flaps in the animation. It's hard to do that while also conveying the meaning of what is said correctly (enough) -- and practically impossible to do that while sounding natural in English too.
As with seemingly all AI these days - it seems to fail with prosody and simply speaks the very next word with zero regard to the nuance or cadence that author intended, or an understanding of any the words being spoken.
When AI achieves the ability to deliver some of Shakespeare’s greatest soliloquies or monologues, then I’ll pay attention.
Prosody is hard. No AI voice I've heard really nails voice generation with fully smooth and human-like cadence and quality, but they're gradually getting better. Airy voices sound a bit tinny and childish, but even so, they're better than many I've heard, including for the elusive "humanness" quality.
Is this intended for anime autodubbing? There's only one voice (Rowan) that is remotely traditional broadcaster-style.
When AI achieves the ability to deliver some of Shakespeare’s greatest soliloquies or monologues, then I’ll pay attention.