Sovereignty & AI ActPublished 29 June 2026· Updated 17 August 20264 min

GDPR and Generative AI: A Practical Guide for SMEs

By Alexandre Saint-Jean

GDPR and Generative AI: A Practical Guide for SMEs

Audio version

Audio version produced by text-to-speech from the article. Our AI charter

View the slides

ChatGPT to draft customer emails, Copilot to analyse a dashboard, Mistral to automate document processing: these tools are entering companies before the legal question has always been asked. Can you put personal data into them? Under what conditions? Here is the practical framework, without unnecessary jargon.

Can you really put customer data into an LLM?

The short answer is: yes, under conditions. GDPR does not ban entrusting personal data to a third party, it regulates it. The real question is not "is this forbidden" but "have I put in place what makes this legal." Three requirements apply at the same time: a valid legal basis, a signed processing contract, and the principle of data minimisation.

The riskiest case is not using the model, it is training it. Some providers reserve the right to use your exchanges to improve their model, unless you explicitly opt out. That re-training option, in nearly every case, requires the consent of the people whose data you have shared.

This question of long-term control is developed further in use AI without losing control of your data, which sets the strategic frame before diving into legal obligations.

GDPR, in Article 6, lists six possible legal bases. For an SME using an LLM on customer data, three come up in practice. None of them automatically covers every use case.

Legitimate interest (Article 6(1)(f)) is often the first basis invoked for internal or B2B uses. It requires a documented balancing test between your business interest and the privacy rights of the people concerned, and that test is hard to pass once the data is rich or the processor is based outside the EU.

Performance of a contract (Article 6(1)(b)) can cover certain uses if the processing is directly necessary to deliver what was ordered. The word "necessary" is interpreted narrowly: analysing a customer's buying habits to refine your catalogue does not qualify.

Consent (Article 6(1)(a)) is required when no other basis holds, or when the data processed is sensitive. It is the safest basis, but also the heaviest to collect and document at scale.

What is a DPA, and why is it mandatory?

The moment you hand personal data to a provider, that provider becomes a processor under GDPR. Article 28 then requires a Data Processing Agreement (DPA), which sets out what the processor can do with your data, under what conditions, and with what security guarantees.

OpenAI, Microsoft and Google all offer a DPA that can be activated from your account. CNIL sets out the mandatory clauses such a contract must include: subject matter, duration, nature, purpose, categories of data, and the processor's obligations. Without a signed, archived DPA, the processing is non-compliant, even if nothing has technically gone wrong. Keep a dated copy in your record of processing activities.

Why isn't choosing a server in the EU enough?

Microsoft Azure Europe, Google Cloud Paris, AWS Frankfurt: many SMEs assume that picking a European data centre settles the sovereignty question. It is a partial protection, not enough for your most sensitive data.

The US Cloud Act, passed in 2018, lets US authorities compel any company subject to US law to hand over data stored anywhere in the world, including in Europe. Microsoft, Google and Amazon are US legal entities. Hosting your data in their European data centres does not shield it from that law.

For real protection, you need a provider governed exclusively by EU law, with no ownership link to a US entity. That is what the French ANSSI's SecNumCloud framework certifies. Providers such as Mistral or OVHcloud meet this requirement: the options and trade-offs are covered in choosing a sovereign LLM for your business.

Anonymisation or pseudonymisation: what's the practical difference?

Pseudonymisation replaces a direct identifier, such as a name or an email, with an alias or an internal code. The data remains personal data under GDPR because re-identification stays possible. Pseudonymisation reduces the risk but does not remove your legal obligations.

Anonymisation makes re-identification technically impossible and irreversible. If it is genuine, the data falls outside GDPR's scope (Recital 26) and you can process it without a legal basis or a DPA. CNIL notes that genuine anonymisation is hard to achieve on datasets rich in profiles and behaviours. A pragmatic rule: pseudonymise before sending, and never send more than the task requires.

Where should you actually start?

Four actions, in this order, cover the essentials for an SME starting from zero.

Map your usage. List the AI tools already in use. For each one: are you processing personal data, and have you signed a DPA? This mapping takes about an hour and often reveals several tools with no active contract.

Sign the missing DPAs. Most providers offer one in a few clicks. Without your signature, you carry the legal liability alone. Keep a dated copy in your record of processing activities.

Minimise by default. Before sending anything to an LLM, ask: is this field actually necessary? Removing an email address or a customer ID costs little and meaningfully reduces exposure.

Log sensitive usage. Keep a record of prompts that involved personal data: who, which tool, what purpose, what date. That traceability is expected during a CNIL audit and helps you catch problems before they become a reportable incident.

Frequently asked questions

Can you put customer names and emails into ChatGPT?
Not without precautions. You first need to identify your legal basis (legitimate interest or contract performance, depending on the case), sign a DPA with OpenAI, and send only the data strictly necessary for the task. The safest option remains pseudonymising or anonymising before sending anything.
What is a DPA and who needs to sign it?
A DPA (Data Processing Agreement) is the processing contract required under GDPR Article 28 whenever a third party processes personal data on your behalf. OpenAI, Microsoft and Google all offer one through their customer portals. Without a signed, retained DPA, any processing is non-compliant, even if technically nothing went wrong.
Does hosting data in the EU protect against the US Cloud Act?
No. The 2018 US Cloud Act lets US authorities compel any company subject to US law to hand over data, including data hosted in the EU. Server residency does not substitute for a provider governed exclusively by EU law.
What is the practical difference between anonymisation and pseudonymisation?
Pseudonymisation replaces an identifier with an alias: the data remains personal data under GDPR, because re-identification stays possible. Anonymisation makes re-identification technically impossible: the data falls outside GDPR's scope. CNIL notes that genuine anonymisation is hard to achieve on rich datasets.

Sources

Get the AI briefing, no commitment

Free · One email a month · Unsubscribe anytime · Your data is never sold

Free first call

Got an AI project in mind?

30 minutes to scope your need and see how to fund it. No commitment.

Working with companies across France, remote.