Cybersecurity & Trust

Where Does Your Data Go? Questions to Ask Any AI Vendor

Most organisations adopting an AI tool never actually ask where the data goes after it leaves their system. The answer is usually available. Almost nobody asks for it before signing.

Written by Gurubalan G.T. · · 4 min read

A single data packet icon following a dotted path through several unlabelled boxes before reaching a final destination, representing the unclear journey of data through an AI vendor's systems.
A single data packet icon following a dotted path through several unlabelled boxes before reaching a final destination, representing the unclear journey of data through an AI vendor's systems.

Adopting an AI tool means sending your organisation's data — sometimes customer data, sometimes proprietary business information — into a system you do not control and usually cannot fully inspect. This is a companion piece to our guide on data security for enterprise software projects.

Why this question is more urgent for AI vendors than for typical SaaS

A conventional SaaS tool stores your data and processes it according to logic you can reason about. An AI system, particularly one built on a third-party foundation model, may pass your data through multiple layers: your vendor's application, the model provider's API, and potentially the model provider's own logging, monitoring or improvement pipelines. Each layer is a separate place your data might be retained, used, or exposed, and each layer may have a different policy.

Major model providers generally distinguish between consumer-facing free-tier usage, which may be used to improve models, and paid enterprise or API usage, which is typically excluded from model training by default — but exact terms, retention periods and opt-out mechanics vary by provider and change over time. Do not assume a policy from one context (a provider's public chat product) applies to a different context (the same provider's enterprise API) without checking the specific terms currently in force for the specific product you are using.

The questions to ask, specifically

Is our data used to train or improve your models, by default or by opt-out? Get the specific answer for the specific product tier you are purchasing, not a general company policy that may not apply to your contract.

If you use a third-party foundation model, whose policy governs our data once it reaches that model? Your vendor's privacy policy and the underlying model provider's policy may be two different documents with two different sets of terms, and your vendor's marketing may only mention the more favourable one.

How long is our data retained, and where? Retention for debugging, logging or abuse-monitoring purposes often has a different (and longer) timeline than the primary processing purpose, and vendors do not always volunteer this distinction unprompted.

Can our data be deleted on request, completely, including from backups and logs? A "yes" without specificity about backups and logs is an incomplete answer.

What jurisdiction is the data processed and stored in? This determines which data protection regime — GDPR, DPDP, US state laws — actually governs your recourse if something goes wrong, and it may not match the jurisdiction your contract is nominally under.

Do you have a SOC 2 report, and if so, Type I or Type II, covering which criteria? See our dedicated piece on what SOC 2 actually proves for how to read the answer once you have it.

What happens to our data if you are acquired, or if you discontinue this product? Vendor continuity risk is a data risk too, and it is rarely addressed in a standard contract unless you ask.

Can you provide an audit trail of what the AI system did with our data and why? This matters for your own compliance obligations under the AI Act, DPDP and similar frameworks, covered in our companion piece on audit trails — if the vendor cannot produce this, you may not be able to meet your own regulatory obligations regardless of how the vendor's internal practices actually work.

Why "we're SOC 2 compliant" is not itself an answer

A vendor citing a compliance certification is answering a different, narrower question than the ones above. SOC 2 evaluates organisational controls; it does not, by itself, tell you whether your specific data is used for model training, how long it is retained, or what jurisdiction it sits in. Treat certifications as a starting point for the specific questions above, not a substitute for asking them.

What to do with the answers

Write them down, attach them to the vendor contract, and revisit them at renewal — vendor data policies change, sometimes without prominent notice, and a policy you verified at signing may not be the policy in force two years later.

What we do

We ask every AI vendor we recommend to clients these exact questions, and we document the answers rather than relying on marketing claims. If you are evaluating an AI vendor and want the data handling questions asked properly before you commit, that is a conversation we are glad to have.


Kaizen Spark Tech designs and delivers software, AI, automation and digital infrastructure for businesses and institutions.

Cybersecurity & TrustAI vendorsdata privacythird-party riskprocurement
Considering a build? Describe the process and we will come back with a scope and a cost range — including if our view is that software is not the right answer. Get a range Message on WhatsApp