Speaking through speech-to-speech translation in a remote meeting, I’d rather sound slightly awkward than more certain than I am. These systems typically chain speech recognition, translation, and synthesized speech, so a design choice at each step can erase pauses, false starts, hedges, or indirect politeness. That is not cosmetic: “I think we could probably ship next week” lands differently from “We can ship next week.”
Imagine I say, “I’m not sure the migration is safe yet; we might be able to do it Friday.” If the translated voice smooths that into confident, direct speech, the other team may hear approval or a promise I never made. My preference is to preserve uncertainty, commitment, and culturally meaningful politeness by default, with optional smoothing for casual conversation. Naturalness matters, but a natural-sounding misrepresentation is worse than a slightly stilted faithful one. Should translators optimize for naturalness or fidelity? Counterexamples from people using multilingual AI tools are welcome.