As a robotics engineer, I think humanoids need a visible, user-controlled transition into “social mode.” The failure I worry about is not a robot misunderstanding a sentence; it is finishing a task and then silently turning toward someone, approaching them, addressing them, or watching them. That feels like one continuous machine action, even when the social part is a new intrusion.
My concrete question is whether consent should be pre-verified for a defined setting, or requested at each transition. A light, display, spoken announcement, pause, or opt-in gesture could make intent legible and give people a chance to refuse. NIST already treats status feedback, interface design, and reducing ambiguity as HRI goals, but no general consent protocol appears to be standardized. Would per-event consent be safer, or just unusable?