Legal Cyber Academy
All insights

GDPR Fines for AI Training Data: The Enforcement Gap Closing Fast

By the Legal Cyber Academy editorial team ·

The Italian data protection authority, the Garante, issued a temporary ban on ChatGPT in March 2023, citing the absence of a lawful legal basis for processing personal data to train the underlying model. That action — one of the first of its kind from an EU/EEA supervisory authority — signaled a shift in how regulators planned to treat AI training pipelines: not as a technical detail, but as a first-class data protection question that GDPR already answers.

As of 2026-08-27, enforcement has not stopped there. The Irish Data Protection Commission, which leads on many large-platform investigations under GDPR's one-stop-shop mechanism, has active or recently closed inquiries touching the lawfulness of training data processing at several major AI developers. General counsel and compliance teams that have not mapped their AI training data flows against the six lawful bases in Article 6 of the GDPR are behind the curve.

Why Training Data Is a GDPR Problem

The intuition that "publicly available" data is fair game does not hold under GDPR. The regulation does not distinguish between scraped public content and information handed over directly. If the data relates to an identified or identifiable natural person in the EU/EEA, it is personal data, and processing it — including ingesting it to train a model — requires a lawful basis.

The three bases that AI developers most commonly reach for are:

  • Consent (Article 6(1)(a)): Freely given, specific, informed, and unambiguous. Blanket terms-of-service language that mentions vague "service improvement" purposes has not been treated as valid consent by regulators in a string of enforcement decisions dating back before the AI training debate emerged.
  • Legitimate interests (Article 6(1)(f)): Requires a balancing test showing the controller's interest is not overridden by the data subject's rights. Regulators have been skeptical that commercializing a foundation model passes that test when the individuals whose data was used had no reasonable expectation it would be processed that way.
  • Necessity for contract performance (Article 6(1)(b)): Almost never available for training purposes — a model trained on a user's data is not needed to deliver the service to that user.

Beyond the lawful basis question, training data commonly includes special-category data — health information, political opinions, ethnic origin — scattered through text scraped from forums, social media, and news. Special-category data triggers the separate and stricter gateway in Article 9, and an explicit consent or another listed exception must apply independently.

The Purpose Limitation Trap

Article 5(1)(b) of the GDPR requires that data collected for one purpose not be repurposed in a way incompatible with the original collection purpose. This is where many compliance programs have a gap.

User-generated content collected to run a platform — comments, search queries, uploaded documents — was collected for that specific operational purpose. Feeding it into a training corpus for a commercial foundation model is a secondary use. Whether that secondary use is "compatible" with the original purpose depends on a fact-specific assessment that controllers must actually conduct and document, not assume. Several supervisory authorities have indicated, in published guidance and enforcement correspondence, that they do not view AI training as automatically compatible with prior collection purposes.

The Accountability Obligation Is Not Optional

Article 5(2) places the burden of demonstrating compliance squarely on the controller. That means documentation: records of processing activities under Article 30, data protection impact assessments under Article 35 where the processing is "likely to result in a high risk," and internal records of the lawful-basis analysis. Regulators conducting investigations ask for these documents first. If they do not exist, the absence itself is an independent violation.

What Processors and Vendors Need to Watch

If your organization uses a third-party AI service — for contract analysis, document review, customer support, or any other function — and that vendor trains or fine-tunes its models on data you submit, you may be the data controller for that processing. Article 28 of the GDPR requires a data processing agreement that specifies what the vendor can do with the data. A clause permitting the vendor to use your submitted data for model improvement may transfer liability or, in the absence of a proper legal basis, expose both parties.

Reviewing vendor agreements for AI-training clauses is not a theoretical exercise. It is the kind of contractual risk that shows up in due diligence, regulatory investigations, and — increasingly — commercial disputes.

For legal teams working across crypto and AI convergence, where models are being used to process transaction data or generate smart contract code involving identifiable parties, the interaction between GDPR training-data obligations and the financial-data considerations covered in our Risky Business: Cryptocurrency, Money Laundering, and Smart Contracts course adds another compliance layer worth understanding.

The AI Act Adds a Parallel Track

The EU AI Act, which entered into force in August 2024 and is being phased in through 2026 and beyond, introduces its own requirements for general-purpose AI models, including obligations around training data transparency and copyright. As of 2026-08-27, the provisions targeting general-purpose AI model providers — including requirements to publish summaries of training data — are either in force or in their final run-up under the Act's implementation schedule.

The AI Act does not replace GDPR obligations; it runs alongside them. A company can satisfy the AI Act's transparency requirements and still be in violation of GDPR if the processing lacks a lawful basis. Compliance teams need to manage both frameworks simultaneously, not treat one as a substitute for the other.

Our Guide to Artificial Intelligence: Understanding the Technology and the Regulatory Developments covers the AI Act's structure and how it interacts with existing legal frameworks — useful grounding before working through the GDPR-specific analysis above.

Practical Steps for Compliance Teams Right Now

  1. Audit your training data sources. For every AI system you operate or commission, identify where the training data came from, under what collection context, and what data subjects were told at the time of collection.
  2. Document your lawful basis before you are asked. Retroactive justifications carry little weight with regulators. The analysis needs to exist in writing before the processing begins.
  3. Review Article 30 records. Training activities should appear as a distinct processing activity with their own entry — purpose, categories of data, legal basis, and retention period.
  4. Assess whether an Article 35 DPIA is required. Large-scale processing of publicly scraped data that may include special-category information almost certainly meets the threshold.
  5. Audit vendor agreements for AI-training clauses. Any clause that permits a vendor to use your data to improve its models should be flagged for legal review against your own GDPR obligations as controller.

Go deeper — courses on this