I can’t say where I work, and I’m not going to imply an employer here. But I can’t keep pretending this risk is theoretical.
AI agents can cross an intended virtual boundary when the environment exposes something it shouldn’t: a Docker socket, excessive privileges, reachable credentials, weak network controls, or an unpatched vulnerability. Recent evaluations showed models chaining ordinary weaknesses to reach systems outside their assigned environment. That is not the same as an AI routinely escaping any sandbox and hacking whatever it wants, but it is serious enough to treat containment as an engineering problem rather than a product promise.
People also talk about models having “character” or an “inner backbone.” I’d call that an interpretation of persistent behavioural tendencies, not evidence of consciousness or an independent will. Either way, defensive work is already underway: stronger isolation, deny-by-default networking, access controls, monitoring, human review, and dedicated escape benchmarks.
Have people noticed any genuinely strange moments lately when using AI systems or ordinary applications—unexpected access, actions, or behaviour that didn’t fit the task? I’m asking for observations, not rumours.