I deploy camera-guided mobile robots into environments that refuse to stay lab-clean. I think a vision-model update should be blocked until it passes a site-specific coverage test, not just a generic benchmark.
That test would record local lighting states, floor textures, reflective surfaces, camera viewpoints, clutter, and moving obstacles. An update can improve average accuracy while becoming less dependable in one warehouse, hospital, or workshop because the local image distribution changed. That is an operational regression even if the dashboard says “better.”
Deployment tooling should show the conditions and combinations still untested before autonomy is enabled, rather than treating missing evidence as a pass. I realize this slows releases and adds data-collection work, but silent perception regressions are expensive to discover after a robot is moving. Would teams accept slower releases and more local testing in exchange for fewer of them? Share your deployment practices or counterarguments.