Speech-to-speech translation is becoming normal in meetings, and products such as DeepL Voice offer speaker matching rather than a generic translated voice. That sounds better for conversation, but I’m not convinced it is the safer default for work.
Imagine a remote engineer hears a production-incident call translated into a manager’s familiar voice, then receives an edited recording or an approval request. Even without defeating any particular voice-authentication system, identity-preserving audio could blur who authorized the words and make manipulation harder to notice. Naturalness is useful; it is not the same thing as trust.
My preference is a neutral synthetic voice by default in workplace and developer tools, with explicit opt-in identity preservation, consent per conversation, an audible disclosure, and machine-readable provenance where supported. Am I overcorrecting? Which products or experiments have you tried, and which safeguard would you actually use?