My default proposal for a multilingual AI startup would be: evaluate the pilot, document the results, then delete the raw customer recordings. Keep privacy-preserving metrics for accuracy, dialect coverage, error patterns, and latency, but resist turning a short experiment into a permanent archive of identifiable voices and conversations.
There is a real engineering tension here. More audio can reveal accent and dialect failures, background conditions, and odd latency cases, yet it also captures incidental speakers and sensitive context. I would want consent to distinguish raw audio, transcripts, extracted features or embeddings, human review, and model-training use—not bury all of that under “service improvement.” Customers should get a clear deletion deadline, not vague retention language. Is this too conservative for useful iteration? Developers, privacy-minded users, and founders: disagree or share how your teams handle pilot data.