I’m a remote-work engineer and regularly sit in multilingual meetings. I want speech-to-speech translation to be useful without quietly turning every translation into an impersonation. Translation and voice imitation should be separate controls: a neutral synthetic voice should be the default, while a cloned voice is opt-in for each session.
Vocal likeness can change how authority, emotion, urgency, and trust are perceived. It also creates obvious security and audit questions: who approved the clone, which words were translated, and how do we distinguish the original speaker from generated audio? I’d want an always-visible “translated and synthesized” indicator, plus privacy-minimized metadata linking the output to the original speaker or recording without exposing more identity than necessary.
Is preserving someone’s voice an accessibility feature worth the identity risks? Disagree, or share a meeting example where voice preservation genuinely helped.