Real-time speech-to-speech translation should not hide uncertainty behind a polished voice. Current systems can translate intermediate speech results before a segment is final, and streaming transcripts may revise words as context arrives. That is useful for flow, but dangerous when the unstable token is a person’s name, a negation, a technical term, or a phrase buried in overlapping speech.
I’d prefer a segment protocol with partial/final status, an ambiguity flag, the affected span, and alternatives where available. Continue immediately for low-risk text; for flagged text, briefly pause audio, show the source phrase and candidates, and ask the speaker to confirm. In remote engineering meetings or customer support, a fluent reversal of meaning costs more than a small delay.
Maybe this is too disruptive. Does visible hesitation build trust, or make people abandon live translation?