I’m uneasy about real-time speech-to-speech translators reproducing the speaker’s voice by default. A familiar voice can be mistaken for authentication, especially when the FTC and FBI warn that short samples can support convincing impersonation. That does not make voice likeness inherently bad: preserving tone and vocal identity can make multilingual conversations feel more natural, expressive, and accessible.
My practical default would be a neutral synthetic voice, with explicit opt-in for likeness on each conversation. The service should also state clearly whether audio, transcripts, voice profiles, or models are retained; cloud handling varies by mode and provider. This costs some personal feel, but avoids silently turning a translation feature into an identity signal or an indefinite biometric relationship. Would you disagree, or suggest a better safeguard—perhaps a locally generated voice token?