DPIA for an internal LLM: when GDPR Article 35 requires a data protection impact assessment for AI
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- Article 35(1) GDPR requires the controller to carry out a data protection impact assessment before processing that is likely to result in a high risk to the rights and freedoms of natural persons, in particular processing using new technologies
- The EDPB’s data protection guide for small business lists nine criteria, among them evaluation or scoring, sensitive or highly personal data, data processed on a large scale and innovative use or applying new technological or organisational solutions, and says processing that meets two of them should in most cases be assessed through a DPIA
- Under Article 35(7) a DPIA contains at least a systematic description of the processing and its purposes, an assessment of necessity and proportionality, an assessment of the risks and the measures envisaged to address them
- For an internal LLM, IT supplies the data flows of prompts, uploaded documents, retrieved passages, answers and logs, where the model runs, the retention of each store, the access model per document and the suppliers with access
- The AI Act asks deployers of high-risk systems to use the provider’s information for their DPIA (Article 26(9)) and lets a fundamental rights impact assessment refer to an existing DPIA (Article 27(4)); for Annex III uses these rules apply from 2 December 2027
Eurokommerz × Vixen.UNO: Private AI/ML Talk to an expert →
DPIA for an internal LLM: when GDPR Article 35 applies
An internal LLM assistant that processes the personal data of employees, customers or other people on a large scale will often meet the criteria that call for a data protection impact assessment (DPIA) under Article 35 of the GDPR, more so when it is used to evaluate people. Article 35(1) requires one before processing that “is likely to result in a high risk to the rights and freedoms of natural persons”, and names processing “using new technologies” as a case in point. The EDPB’s data protection guide for small business lists nine criteria and says that processing meeting two of them should in most cases be assessed through a DPIA. A company-wide assistant with retrieval over HR, customer or case files can meet several of them.
Most of the facts a DPIA rests on come from IT: what flows where, where the model runs, how long prompts and logs are kept, who can read which document and which suppliers can reach the data. This guide lists those inputs for a platform with 500 to 2,000 users. Whether a DPIA is required, and what it concludes, is an assessment for the controller, its data protection officer and its legal department.
What GDPR Article 35 requires
Article 35(1) obliges the controller to carry out the assessment “prior to the processing”, and adds that “a single assessment may address a set of similar processing operations that present similar high risks.” Under Article 35(2) the controller “shall seek the advice of the data protection officer, where designated”. Article 35(3) names three cases in which a DPIA is required in particular. Point (a) covers “a systematic and extensive evaluation of personal aspects relating to natural persons which is based on automated processing, including profiling”, where decisions based on it produce legal or similarly significant effects. Point (b) covers “processing on a large scale of special categories of data referred to in Article 9(1)” or of data on criminal convictions, and point (c) covers large-scale monitoring of a publicly accessible area.
Article 35(7) sets the minimum content: “a systematic description of the envisaged processing operations and the purposes of the processing”, “an assessment of the necessity and proportionality”, “an assessment of the risks to the rights and freedoms of data subjects”, and “the measures envisaged to address the risks, including safeguards, security measures and mechanisms”. Article 35(9) asks the controller, where appropriate, to “seek the views of data subjects or their representatives”. Article 35(11) requires a review, where necessary, “at least when there is a change of the risk represented by processing operations.”
EDPB criteria and the EDPB opinion on AI models
The EDPB’s data protection guide for small business lists the nine criteria: evaluation or scoring; automated decision making with legal or similar significant effect; systematic monitoring; sensitive data or data of a highly personal nature; data processed on a large scale; matching or combining datasets; data concerning vulnerable data subjects; innovative use or applying new technological or organisational solutions; and processing that prevents individuals from exercising a right or using a service or a contract. It states that “in most cases, processing operations meeting two of the following criteria should be assessed through a DPIA” and links the Article 29 Working Party’s Guidelines on DPIA (WP248 rev.01). The EDPB endorsed them on 25 May 2018, in the version revised and adopted on 4 October 2017.
In Opinion 28/2024, adopted on 17 December 2024, the EDPB calls DPIAs “an important element of accountability” where processing in the context of AI models is likely to result in a high risk, in a list of provisions the opinion does not analyse. It considers that “AI models trained on personal data cannot, in all cases, be considered anonymous”, so weights fine-tuned on employee or customer records are one more store in the data flow. Its executive summary says that supervisory authorities should take into account whether the controller deploying a model “conducted an appropriate assessment” to ascertain that the model was not developed by unlawfully processing personal data. The model record in the section on suppliers below supports that assessment.
Large-scale processing of health data falls under Article 35(3)(b); our guide to a private LLM for hospitals covers that case.
What IT supplies for each part of the DPIA
On an LLM platform, the prompt and any uploaded file reach the gateway and the model server. Retrieval adds passages from the vector index, which stores text chunks and their embeddings. The answer goes back to the front end, which keeps the conversation in its database. The gateway log may hold prompts and answers in full, and backups copy all of it. The KV cache in GPU memory holds the context while a request runs, longer with prefix caching.
| DPIA ELEMENT | GDPR REFERENCE | INPUT FROM IT |
|---|---|---|
| Description of processing | Art. 35(7)(a) | data flow diagram with every store, categories of data subjects, model and engine versions |
| Purposes | Art. 35(7)(a) | use-case register with purpose, owner and user directory group |
| Necessity, proportionality | Art. 35(7)(b), 5(1)(c) | connectors and index scope per use case, log fields, chosen retention |
| Risks to data subjects | Art. 35(7)(c) | threat model: retrieval leaks, prompt injection, log access, external APIs |
| Measures | Art. 35(7)(d), 32 | permission checks at retrieval, encryption, log access rules, deletion jobs, restore tests |
| Recipients, transfers | Art. 30(1)(d), (e) | suppliers with access, hosting location, any external API, remote support |
| Erasure | Art. 5(1)(e), 17(1) | retention per store and deletion by user and by source document |
Regulation (EU) 2016/679, Articles 5, 17, 30, 32 and 35, on eur-lex.europa.eu as of October 2026; the right-hand column is our example of the inputs.
Article 30(1)(f) asks the record of processing to give, where possible, “the envisaged time limits for erasure of the different categories of data”, so set a period for each store: chat history, gateway log, vector index and backups. An erasure request under Article 17 can concern the data in all of them, so the index should support deletion by source document and the logs deletion by user ID. Deleted data leave backups only when the backups expire. Our guide to ISO/IEC 42001 and AI governance for a private LLM describes the use-case register and the query log that feed the first rows.
Our Private AI/ML service builds the platform with a query log and data and permissions management, so security and legal see who accesses what. Send us the use cases and the data each one would reach through the form below.
Worked example: an assistant for 1,500 employees
Take the platform from our governance guide: 1,500 staff, two GPU servers on-premise, three open-weight models behind one gateway and six approved use cases. The table shows five of them (not the code assistant) and one proposed use, with the EDPB criteria to check for each.
| USE CASE | PERSONAL DATA IN IT | CRITERIA TO CHECK |
|---|---|---|
| Drafting and translation | whatever staff paste in | large scale; innovative use |
| Policy search with RAG | authors and names in policies | innovative use |
| Contract summaries, legal | signatories, counterparties | innovative use; large scale by volume |
| Ticket classification | names and contact data of requesters | large scale; matching datasets |
| HR FAQ for employees | questions about own leave, pay, health | sensitive or highly personal data |
| Screening applications (new) | applicants’ CVs and letters | evaluation or scoring |
Criteria from the EDPB’s data protection guide for small business; use cases from our AI governance guide; the mapping is our example, not a legal classification.
Article 35(1) allows one assessment for a set of similar operations, so the first three uses could share one if the controller, advised by the DPO, finds that they present similar high risks. In the HR FAQ employees describe their own situation, so a shorter retention period and fewer log readers for that use are measures to weigh. Screening applications would be a new purpose with its own assessment, and it is also listed in Annex III, point 4(a), of the AI Act as “to analyse and filter job applications, and to evaluate candidates”.
Risks and measures on the LLM platform
Article 32(1) lists measures “as appropriate”, among them “the pseudonymisation and encryption of personal data” and regular testing of their effectiveness.
| RISK | PLATFORM MEASURE | GDPR REFERENCE |
|---|---|---|
| Restricted file in an answer | permission filter inside the retrieval query, identity from sign-in | Art. 5(1)(f), 32(1)(b) |
| Planted instructions | list of writers per source, approval before tool calls, injection tests | Art. 32(1)(d) |
| Logs read too widely | log store open to named roles only, access to it logged | Art. 32(1)(b) |
| Prompts kept too long | retention per store, deletion jobs, backup expiry | Art. 5(1)(e) |
| Data sent to a public API | external APIs off by default, enabled per use, visible in the log | Art. 28, 44 |
| Wrong statement on a person | cited sources, human review where output concerns individuals | Art. 5(1)(d) |
| Loss of index or logs | backup and a tested restore | Art. 32(1)(c) |
Regulation (EU) 2016/679, Articles 5, 32 and 44, on eur-lex.europa.eu as of October 2026; measures are our examples.
Our guide to prompt injection and LLM security covers the tests and the review before go-live behind these measures.
Engineering by our partner Vixen.UNO covers protection against prompt injection, access rights and logging of queries and answers. Describe your sources, users and current logging in the form below.
Suppliers, hosting and the model’s origin
When an open-weight model runs on the company’s own servers, prompts do not reach the model’s publisher. The DPIA lists the suppliers that can reach the data, such as a hosting provider, an engineering partner with administrative access and any public API enabled for some tasks. Under Article 28(1) the controller “shall use only processors providing sufficient guarantees”, and Article 28(3)(f) has the processor assist the controller with “the obligations pursuant to Articles 32 to 36”, which include the DPIA. What that contract must say is the subject of our guide to GDPR cloud hosting, the Article 28 DPA and subprocessors.
For each model, keep a record that supports the deploying controller’s assessment: the publisher and where the weights came from, the checksum checked on import, the model card and licence, and which version answered which use case.
How the AI Act deployer duties relate to the DPIA
For high-risk AI systems, Article 26(9) of the AI Act asks deployers to use the information the provider supplies under Article 13 “to comply with their obligation to carry out a data protection impact assessment under Article 35”. Article 26(6) has them keep automatically generated logs, to the extent they control them, for a period “of at least six months, unless provided otherwise in applicable Union or national law, in particular in Union law on the protection of personal data”, so the retention period in the DPIA and the AI Act log period have to be set together.
Article 27(1) requires a fundamental rights impact assessment before a high-risk system under Article 6(2) is deployed, with critical infrastructure (Annex III, point 2) excepted. It applies to deployers that are bodies governed by public law or private entities providing public services, and to deployers of the credit scoring and life and health insurance systems in points 5(b) and (c) of Annex III. Article 27(4) lets such an assessment “include cross-references to the relevant sections of that data protection impact assessment”. Under Article 113 as amended by Regulation (EU) 2026/1744, these rules apply to Annex III uses from 2 December 2027. Drafting, search and summaries are not Annex III purposes; our guide to EU AI Act deployer obligations sets out what applies to them today.
Prior consultation, review and sign-off
Article 36(1) requires the controller to consult the supervisory authority “where a data protection impact assessment under Article 35 indicates that the processing would result in a high risk in the absence of measures taken by the controller to mitigate the risk.” Article 39(1)(c) gives the DPO the task to advise on the DPIA where requested “and monitor its performance”. IT owns the technical annex and keeps it current through the change process.
- Each change request names the use case, model, connector or retention setting it touches.
- The platform owner records whether data categories, recipients, model or storage change.
- If they do, the controller reviews the DPIA with the DPO’s advice, since Article 35(11) foresees a review when the risk changes.
- The DPIA and the use-case register are updated before the change goes live.
What we do
Our Private AI/ML service builds private LLM platforms on-premise or on dedicated hardware in a Tier-3 data centre in Lithuania, with engineering by our partner Vixen.UNO. It includes protection against prompt injection, data and permissions management, and logging of queries and answers. Nothing goes to public services unless you explicitly enable it. We deliver the technical part and train your team to run the platform. A data processing agreement is provided on request, with subprocessors named in the contract, as our security and compliance page sets out.
FAQ
Does an internal LLM need a DPIA?
What is a DPIA under GDPR Article 35?
What must a DPIA for generative AI contain?
Do I need a DPIA for a chatbot?
Who carries out a DPIA for an AI system?
How does a DPIA relate to the AI Act fundamental rights impact assessment?
Send us your use cases, the number of users, the data each use would reach and where the platform should run. We reply within one business day and arrange a first call, in which you get 2 to 3 possible solution scenarios for the platform and its technical controls. The first call is free of charge.
Talk to an expertWe reply within one business day