Where the data goes
Every model call sends something somewhere. For many organisations that is entirely acceptable and for some it is not, and the distinction is usually contractual rather than philosophical. Customer agreements may restrict where data is processed. A data processing agreement may name permitted subprocessors. A sector regulator may have a position on cross border processing. These constraints exist before any technical decision and they narrow the options in ways that are much cheaper to discover early.
We work through it per use case rather than as a blanket policy, because the answers differ enormously. Summarising public marketing copy raises almost nothing. Reading employee grievances or customer contracts raises a great deal. For each, the record states what data is sent, to which provider and region, how long they retain it, whether inputs may be used for model improvement, and what the fallback is if the answer becomes unacceptable and you need to move.
Reducing what is sent is usually the most effective control available. Redacting identifiers before a call, sending an extract rather than a whole document, retrieving the specific passage rather than the whole file, and keeping deterministic work local all reduce exposure without reducing usefulness. Where the requirement is that nothing leaves your environment, that is a legitimate constraint with real consequences for capability and cost, and it should be stated as a decision rather than assumed away.
- A written data handling position per use case rather than one policy covering everything
- Provider region, retention period and training use recorded and checked against your agreements
- Personal and commercially sensitive data minimised or redacted before any external call
- An exit position stated per use case, so a change in provider terms is not a crisis
- Local or private deployment considered where the constraint genuinely requires it