A data protection impact assessment is very likely required for a generative AI project that processes personal data, because GDPR Article 35 mandates a DPIA whenever processing is likely to result in a high risk to individuals, and European data protection authorities have generally treated large-scale use of novel AI technology as meeting that bar by default. Specific triggers that push a generative AI project toward a mandatory DPIA include large-scale processing of personal data, systematic profiling or scoring of individuals, use of special category data such as health records, and the general novelty and opacity of the technology itself, all of which the EDPB has flagged as high-risk indicators. A DPIA for this kind of project should assess data flows into the model, the risk of the model memorizing or leaking training or prompt data, retention periods, and the mitigations in place such as anonymization or access controls. Skipping a DPIA when one is legally required is itself an enforceable violation, separate from any downstream data protection issue. Smaller projects using anonymized or synthetic data with no special category information may fall below the threshold, but that determination should be documented rather than assumed. Nanobase AI works with legal and privacy teams to scope and document DPIAs for generative AI projects before development begins.
Why regulators default to "yes" for generative AI
GDPR Article 35 requires a data protection impact assessment whenever processing is likely to result in a high risk to individuals, and European data protection authorities have generally treated large-scale generative AI processing as clearing that bar by default rather than as a case-by-case judgment call. This is different from how DPIA requirements are sometimes applied to more established processing activities, where a company can often reason its way to "no DPIA needed." For a generative AI project processing personal data at any meaningful scale, the safer starting assumption is that a DPIA is required, and the analysis should focus on scoping it correctly rather than debating whether it applies.
The specific triggers to check
| Trigger | Present in most generative AI projects? |
|---|---|
| Large-scale processing of personal data | Often yes, if customer or employee data flows through prompts |
| Systematic monitoring or profiling | Yes, if outputs are used to score, rank, or predict behavior |
| Processing of special category data (health, biometric, etc.) | Depends on use case; common in healthcare and HR applications |
| Use of new or innovative technology | Generally yes, since large language models are treated as novel processing |
| Automated decision-making with legal or similarly significant effect | Yes, if the AI's output directly determines an outcome without human review |
A project hitting even two of these triggers is a strong candidate for a mandatory DPIA, and generative AI projects frequently hit three or more simultaneously. Checking the triggers individually, rather than debating the project's overall risk in the abstract, is what turns this into a fast yes-or-no determination.
Running the assessment: the parts that are different for LLMs
- Describe the processing systematically, including every data flow from prompt to output, not just the "AI feature" at a high level.
- Assess necessity and proportionality, specifically whether the task requires personal data at all, or whether redaction or synthetic data could achieve the same result.
- Identify risks specific to generative AI: hallucinated output presented as fact, potential memorization and leakage of training or fine-tuning data, and prompt injection causing unintended data exposure.
- Define mitigations for each identified risk, such as output review steps, retention limits, and access controls on the underlying model and logs.
- Consult the data protection officer, and consult the supervisory authority in the rare case where residual risk remains high after mitigations.
The risk-identification step is where generic DPIA templates fall short for LLM projects, since hallucination and memorization are not risks a traditional software DPIA template anticipates.
Common scoping mistakes
Two mistakes recur often. Treating the DPIA as a one-time document rather than a living assessment updated when the model, vendor, or data flow changes is the most common, especially since LLM applications iterate faster than typical enterprise software. The second is scoping the DPIA around the AI vendor's model card instead of the company's own specific deployment, which misses risks introduced by the company's own prompt construction, retrieval logic, or logging practices. A DPIA scoped to the vendor's model rather than the company's actual deployment misses most of the risk the assessment is supposed to catch. This is general information, not legal advice, and the DPIA process should involve the organization's data protection officer or qualified counsel.
Frequently asked questions
Who is responsible for completing the DPIA, the vendor or our company?
The data controller, which is almost always the company deploying the generative AI feature for its own purposes, is responsible for the DPIA, even when the underlying model is provided by a third party like OpenAI or Anthropic.
Can we reuse one DPIA across multiple generative AI features?
Only if the features share substantially the same processing, data types, and risk profile. Distinct features, such as an internal coding assistant and a customer-facing chatbot, typically need separate assessments since their data flows and risks differ meaningfully.
What happens if we skip a required DPIA?
Failing to carry out a required DPIA is itself a GDPR violation independent of whether any actual harm occurred, and it removes the documented risk analysis that would otherwise support a defense if a regulator investigates the processing later.
Does a DPIA replace the EU AI Act's risk classification?
No, they serve different purposes and are often run in parallel: the DPIA assesses data protection risk to individuals, while AI Act classification assesses the system's risk tier under a separate framework, as described in our checklist for EU AI Act, GDPR, and KVKK compliance.
How Nanobase AI helps
Nanobase AI, an enterprise AI engineering company based in Silicon Valley, supports compliance and engineering teams in scoping and running DPIAs specific to generative AI projects, identifying the hallucination, memorization, and injection risks that generic templates miss. This is part of our AI security and compliance work, which extends into the broader GDPR compliance steps covered in how to make an LLM application GDPR compliant.
Ready to discuss your project? Contact Nanobase AI or email hello@bumu.tech.