AI in the crosshairs of supervisory authorities: a checklist to prepare for audits

16 min

Supervisory authorities are increasingly focusing on the use of artificial intelligence. While the AI Act sets the framework at European level, data protection authorities are already examining whether AI systems are compatible with the GDPR. Companies therefore face a twofold challenge: they need to be able to use innovative AI applications without coming into conflict with data protection requirements, while at the same time being prepared for possible audits. We explain what the authorities look for, where typical conflicts with AI arise and how an audit actually works in practice. Finally, you will receive a checklist to help you prepare your own organisation for an audit by the authorities.

Schedule a no-obligation initial consultation now

What the authorities actually examine

Data protection supervisory authorities operate within the area of competence assigned to them. Accordingly, they examine compliance with the principles of the GDPR, i.e. the GDPR-compliant processing of personal data and the accountability that must be demonstrated to the authority in this respect.

The challenge in the AI context is that most of the GDPR principles are, in part, at odds with the use of AI. It is important to show that you have taken the AI-specific risks into account in the data protection assessment and have properly addressed and resolved apparent contradictions.

Disagreement between supervisory authorities

To gain initial guidance on the specific requirements of the data protection supervisory authorities, it is worth looking at the current publications of the German state data protection authorities, the German Data Protection Conference (DSK) and the European Data Protection Board (EDPB) on the interface between data protection and AI applications.

Since there is not yet any case law from the Court of Justice of the European Union (CJEU) on the specific requirements, these papers do not claim to offer reliable solutions. Instead, they merely make suggestions and provide a basis for orientation.

However, the authorities' papers can serve as a guide to what supervisory authorities expect. It should be noted, though, that even the supervisory authorities do not always agree on data protection issues.

One example of this is the Hamburg data protection authority's discussion paper on large language models (LLMs), which caused a stir in the data protection world. The authority takes the view that no personal data is stored in LLMs.

This contrasts with the position of the Baden-Württemberg supervisory authority and would mean that data subject rights would be limited to the input and output of an AI system and could not extend to the model itself. Against this background, it is important for advisers to take a position and develop a feel for the best solution.

Reviewing the various documents published by the authorities therefore gives companies an overview of what the supervisory authorities require, but not HOW to deal with it. It is striking that this jungle of documents leaves companies that source AI from third parties (AIaaS/service providers) out in the cold and makes only very vague statements about this scenario.

It is therefore important to take a systematic approach and keep the AI lifecycle in mind, e.g. based on the MLOps model. This means that each "life phase", from selecting and training the AI model, through testing and validation, to using the AI system, must be assessed separately from a data protection perspective.

Schedule your initial consultation

Describe your situation to us in a no-obligation phone call, and our lawyers will work with you to find the best solution.

Schedule consultation

Where AI and the GDPR collide

Transparency vs. black box

Data subject rights require traceability, while modern AI models often remain opaque.

The transparency requirements, particularly with regard to data subject rights such as the right of access and the right to information, conflict with the fact that, with the current GenAI models, AI decision-making processes cannot be fully traced in detail (the black box phenomenon).

As a result, the mandatory information under Article 13 GDPR or access requests may not always be easy to fulfil, because the company does not have sufficient information, because doing so would mean disclosing trade secrets, or because the information does not meet the required level of comprehensibility.

Complete and transparent information within the meaning of Articles 13 and 14 GDPR about data processing carried out by means of AI is therefore often hardly possible. The "explainable AI" (XAI) approach attempts to solve this problem by using various methods to shed light on the matter. One of these approaches is post-hoc methods such as SHAP (Shapley Additive Explanations).

This method seeks to determine and explain the influence that each input variable has on the result. This allows conclusions to be drawn about the AI's decision-making process, which improves traceability overall.

Another method is known as "counterfactual explanations". Here, the AI is given an input, and testing is then used to determine which changes to the input would have led the AI to a significantly different result. This narrows down the decisive parameters of the decision and makes AI decisions traceable.

Regardless of the method chosen, the data flows associated with the AI applications must in any case be mapped phase by phase in the record of processing activities, so that at least the internal documentation is in place as a basis for meeting the transparency obligations.

Newsletter

For your Inbox

Current updates and important information on topics such as data law, information security, technology, artificial intelligence, and much more. (only in German)

Please calculate 3 plus 4.

By clicking on the button, you consent to receiving our newsletter and to the aggregated usage analysis (opening rate and link clicks). You can revoke your consent at any time, e.g. via the unsubscribe link in the newsletter. More information: Privacy policy.

Purpose limitation vs. training data

Historical data may not simply be reused at will, unless it has been effectively anonymised or is covered by AI regulatory sandboxes.

The principle of purpose limitation laid down in Article 5(1)(b) GDPR states that collected data may only be processed for the purpose for which it was collected. In practice, however, it is precisely historical data that companies want to use for AI training. Under the principle of purpose limitation, this data cannot therefore readily be used for AI training outside its original purpose.

One possible solution is to use anonymised or synthetic data. Anonymised data is data that was originally personal but has been altered in such a way that it no longer allows any conclusions to be drawn about the individual. If anonymisation is effective, this data falls outside the scope of data protection law and can therefore be used freely for AI training.

With anonymised data, however, the problem arises, particularly for small and medium-sized enterprises, that robust anonymisation can involve considerable effort. There is also a risk that anonymised data used to train an AI and later reproduced by it could, in combination with external datasets held by a third party, allow the data subjects to be re-identified after the fact.

Some legal commentators take a very strict view, according to which data is only anonymised if re-identification is not possible under any circumstances, including with future methods. If you follow this view, anonymised data cannot in practice be used for AI training without incurring a liability risk. In addition, there is a risk that anonymised data will be unsuitable for AI training because the data point relevant to the training had to be anonymised.

Another option for training is the use of synthetic data. This is data that has been generated entirely artificially. It is modelled on real personal data but does not relate to real people. Under certain circumstances, synthetic data can produce comparable learning effects in AI training. However, there is a risk that training exclusively with synthetic data creates an echo chamber.

To illustrate this, imagine an AI system that corrects its own review results and, in the next step, reads and evaluates the correction of the correction. In the best case, the AI system recognises its own errors because it has already received sufficient training material and, on that basis, gradually removes errors until the result is error-free. However, there is also a risk that the AI system does not recognise the errors, but instead identifies the faulty part as correct and repeats the same error on a larger scale.

However, Article 59(1) of the AI Act creates an exception to the principle of purpose limitation. It provides a legal basis under data protection law that permits the further processing of personal data for the development of AI systems in an AI regulatory sandbox, subject to numerous strict conditions.

AI regulatory sandboxes are experimental spaces in which companies and other stakeholders can test AI systems in a protected setting but under real-world conditions before they are placed on the market. The ability to develop and test AI systems under simplified rules creates room to try out innovations. The parties involved should be able to identify any risks and opportunities at an early stage and make decisions on that basis.

In practical terms, such sandboxes are based on so-called experimentation clauses, which allow the competent authority to grant stakeholders exemptions from requirements and prohibitions. Sandboxes therefore serve to strike a balance between state requirements and prohibitions on the one hand and the interests of the parties involved and of the public in progress and further development on the other.

Storage limitation vs. big data

AI needs large volumes of data, while the GDPR requires timely deletion. Only consistent Data Governance can provide clarity here.

AI models require extensive datasets that remain available over the long term. The GDPR, however, requires personal data to be accurate and up to date and to be deleted as soon as it is no longer needed for its purpose.

It is highly doubtful whether personal data in big data training sets can be checked for accuracy. Correcting or deleting personal data that has already been processed is also virtually impossible. To resolve the conflict between the data protection principles of accuracy and storage limitation in AI training, a company should embed consistent Data Governance from the outset.

This first requires data sources to be analysed and documented in a structured manner, and the permissible further processing purposes to be reviewed and defined. Changes of purpose and processing based on legitimate interests, including deletion periods, must be documented.

Technical pipelines then ensure that only reviewed, standardised, de-duplicated and versioned datasets from approved data sources make it into training.

Integrity and confidentiality vs. AI service providers

Involving US providers brings additional risks, especially if there is no DPF certification.

When data is processed in AI systems, the integrity and confidentiality of the data processed must be ensured (Article 32 GDPR). The AI Act sets out separate requirements for providers of high-risk AI systems. However, challenges arise when using service providers outside Europe, as not every country offers a level of protection equivalent to that of the EU.

Many AIaaS providers are based in the US, and not all of them are certified under the Data Privacy Framework (DPF), which means they are not covered by the adequacy decision between the EU and the US. It must therefore be checked whether the US company that is part of the individual AI value chain and processes personal data in that context is DPF-certified. If this is not the case, other appropriate safeguards must be provided in accordance with Article 46 GDPR.

It is therefore advisable to work with EU-based AI service providers or, alternatively, with US companies where no data is transferred. AI models from US providers can also be used in a fixed form without ongoing data flows to the model provider. It is also necessary to clarify responsibilities with regard to the AI application and the underlying processing activities, in particular whether there is joint controllership or whether it is a case of processing on behalf of a controller, which may require data processing agreements pursuant to Article 28 GDPR to be concluded.

How an AI audit actually works

Although the course of an audit by the authorities is always shaped by the individual case and the current priorities of the supervisory authority concerned, certain basic patterns can be identified.

The focus is always on the authority's competence under data protection law: the only question examined is whether the requirements of the GDPR are being met in connection with the specific use of AI. Companies should therefore be prepared for the fact that, although individual areas of focus are set, the audit steps follow a recurring pattern, from the initial contact through to the request for detailed evidence.

1. Preparation: hearing questionnaire and questions

The first step usually consists of a hearing questionnaire and, where applicable, coordination meetings to schedule the audit, determine the documents required, and agree on access to relevant systems and the participants on both sides. Audit appointments can take the form of pure document reviews or be combined with on-site visits.

In certain constellations, it may be essential to request access to the case file in order to understand what triggered the audit. We have accompanied audits that were announced at very short notice (2 to 3 weeks).

2. From the briefing to the system presentation

The on-site audit usually begins with a briefing on the company's rights and obligations. It is then advisable to present the company and its products or services, as well as the subject matter of the audit from the perspective of the controller. As expected, the meeting then involves a detailed explanation of the individual AI systems and the overall AI architecture.

This covers the specific purposes and objectives, the technical functions of the AI system and its operational use. It should be possible to illustrate the input and the variance of the output in detail for each system. Error risks or the availability risk that may arise from the failure of external service providers can also play a central role in the audit.

3. Detailed questions from the authority on AI training, data and safeguards

The audit will pay particular attention to the use of data for AI training, from both a technical and a legal (data protection) perspective. Companies should be able to explain in detail not only the individual components used, in particular the AI models, but also the reasons for choosing the model, the licences and the tests carried out.

Questions can also be expected about the specific use of data, the categories of data concerned, the scope of data processing and the various sources used. Standardisation and anonymisation procedures should be both described in the abstract and visibly implemented in the system in practice. If synthetic data is used, it should be possible to explain how it was generated. You should expect the authorities to take an interest in how data subject rights are safeguarded in all phases of the AI lifecycle.

Particular attention will also be paid to how AI data processing is explained in the privacy notices for data subjects. Care should be taken to ensure that the explanation reaches the respective target group of data subjects at the right time, is comprehensible in content and suitably concise, and that changes to AI systems are always accompanied by an update to the privacy information under Article 13 GDPR. Versioned privacy information should be kept as evidence.

4. Follow-up: which documents the authority requests

Following an initial audit meeting, you should expect to be asked for further documents. In addition to the presentations shown and detailed technical information on the AI systems and their technical and organisational measures, you should also expect requests for the documentation of the balancing of interests in line with current EDPB standards and the records of processing activities under Article 30 GDPR.

Process descriptions for logging procedures, anonymisation and deletion concepts, access control concepts for databases used for AI training and, finally, the documentation of the data protection impact assessment (DPIA), including the risk analysis, should also be readily available.

Our conclusion: how to make your AI audit-proof

In conclusion, data protection authorities have recognised the productive use of AI in day-to-day business and provide information on it, even though some of this information is contradictory. The authorities have begun examining the use of AI from a data protection perspective, and every company should maintain its mandatory documentation as part of its accountability obligation so that it can respond to queries from the authorities.

Ultimately, it is essential for every company to draw up an appropriate, individual concept for its own use of AI and to document its data protection-compliant use of AI in detail. This is fundamental to being able to act with confidence in a later audit by the authorities without being put on the spot.

Checklist: are you prepared for an AI audit?

1. Documentation of the AI system

  • What: Describe the purpose, function, architecture and area of use of each AI system.
  • Why it matters: Authorities want to understand which data flows exist and which business processes are affected.
  • Practical tip: Create a central "AI dossier" that documents all systems with a brief description of their function, responsibilities and interfaces.

2. Transparency and traceability

  • What: Make sure you can explain the AI system's decisions (e.g. with explainable AI methods such as SHAP or counterfactuals).
  • Why it matters: Transparency is at the core of data subject rights (Articles 13 and 15 GDPR).
  • Practical tip: Have simple explanations ready ("management summary") as well as separate technical documentation for the authority's specialist questions.

3. Data strategy and training data

  • What: Check which data was used for training: legacy/existing data, anonymised or synthetic data.
  • Why it matters: Purpose limitation (Article 5 GDPR) restricts further use. Authorities scrutinise this area particularly closely.
  • Practical tip: Create an overview of where your training data comes from, how it was prepared and whether legal bases (e.g. AI regulatory sandboxes, Article 59 AI Act) were used.

4. Data Governance and deletion concepts

  • What: Define clear processes for data quality, correction and deletion.
  • Why it matters: Authorities expect storage limitation and accuracy to be observed in AI systems too.
  • Practical tip: Rely on automated data pipelines that only allow reviewed and de-duplicated datasets into training. Document deletion periods.

5. Contracts and service providers

  • What: Check whether your AI provider is based in the EU or whether a third-country transfer takes place.
  • Why it matters: Using US providers without DPF certification requires additional safeguards (Article 46 GDPR).
  • Practical tip: Conclude clear data processing agreements pursuant to Article 28 GDPR and document responsibilities, especially in complex AI value chains.

6. Privacy notices and data subject rights

  • What: Update your privacy notices regularly and adapt them to your use of AI.
  • Why it matters: Authorities check whether data subjects can exercise their rights effectively.
  • Practical tip: Keep a versioned collection of all privacy notices. Make sure the texts are understandable for laypeople and legally robust at the same time.

7. Technical and organisational measures (TOMs)

  • What: Protect data integrity and confidentiality (e.g. encryption, access control concepts).
  • Why it matters: Authorities expect evidence under Article 32 GDPR.
  • Practical tip: Document not only existing TOMs, but also regular testing, training and contingency plans.

8. Data protection impact assessment (DPIA) and risk analysis

  • What: Carry out a DPIA for AI systems with increased risk.
  • Why it matters: The use of AI in sensitive areas (e.g. health, employee data) always requires a DPIA.
  • Practical tip: Supplement the DPIA with AI-specific questions (bias, discrimination, black box effects). Involve IT, specialist departments and, where appropriate, external expertise.

9. Internal responsibilities and roles

  • What: Clarify who in the company is responsible for AI systems (data protection officer, AI officer, head of IT).
  • Why it matters: Authorities require clear points of contact and a clear allocation of responsibilities.
  • Practical tip: Set up an interdisciplinary "AI compliance team" made up of legal, IT and business, which documents and prepares on an ongoing basis.

10. Audit simulation

  • What: Run an internal test of how an audit by the authorities would proceed.
  • Why it matters: This allows you to identify gaps and practise your lines of argument.
  • Practical tip: Simulate a scenario with a hearing questionnaire, a meeting with the authority and the provision of evidence. Have specialist departments present their systems.

Schedule your initial consultation

Describe your situation to us in a no-obligation phone call, and our lawyers will work with you to find the best solution.

Schedule consultation