CNIL Recommendations on Applying the GDPR to AI System Development
- April 26, 2026
- Posted by: rob
- Categories: AI & Technology Law, Data Privacy & Cybersecurity
On January 5, 2026, France’s data protection authority, the CNIL, published its first comprehensive recommendations on the application of the General Data Protection Regulation (GDPR) to the development of artificial intelligence systems. These recommendations address a long-standing tension in European privacy law: how to reconcile the GDPR’s stringent data protection requirements with the data-intensive realities of modern AI development.
CNIL is clear that training datasets frequently contain “personal data” within the meaning of the GDPR — information relating to identified or identifiable natural persons — and that the use of such data carries legal obligations that cannot be set aside merely because the processing occurs during a development or research phase.
The CNIL’s recommendations are designed to guide designers and developers through the key compliance checkpoints from initial system design through to dataset creation and model training. Critically, the recommendations explicitly address the dual regulatory environment created by the intersection of the GDPR and the EU Artificial Intelligence Act (AI Act), which was adopted in the summer of 2024. Where personal data is used for AI development, both instruments apply simultaneously, and the CNIL’s framework has been drawn up to supplement them.
I. Scope and Applicability
Which AI Systems Are Covered?
The CNIL’s recommendations apply to any AI system whose development involves the processing of personal data. The three categories expressly within scope are: (1) systems based on machine learning; (2) general purpose AI (GPAI) systems whose operational use is defined from the development phase and which can be used for various applications; and (3) systems for which learning is done “once and for all” or continuously, including through the use of real-time usage data for ongoing improvement.
Practitioners should note that the recommendations are explicitly limited to the development phase of AI systems — that is, all steps prior to deployment in production, including system design, dataset creation, and model training. The CNIL has separately indicated that recommendations addressing the deployment phase are forthcoming.
Relationship to the EU AI Act
The recommendations take the EU AI Act fully into account. The CNIL’s position is straightforward: where personal data is used in the development of an AI system, both the GDPR and the AI Act apply concurrently. The CNIL’s framework is designed to supplement the AI Act in a manner that is consistent with data protection principles, rather than creating a separate or competing compliance regime. For high-risk AI systems within the meaning of the AI Act, the CNIL considers that a Data Protection Impact Assessment (DPIA) is in principle required under the GDPR, and notes that the DPIA may be based on the technical documentation already mandated by the AI Act, provided it incorporates all elements required by Article 35 of the GDPR.
II. The Eleven-Step Compliance Framework
Step 1 Define an Objective (Purpose) for the AI System
The Principle
The foundational obligation under the GDPR’s purpose limitation principle is that any AI system exploiting personal data must be developed with a well-defined objective. This objective must be determined, or at least established, as soon as the project is defined. It must also be explicit (known and understandable) and legitimate (compatible with the organisation’s tasks and lawful under applicable law). Crucially, the objective frames and limits the personal data that can be used for training, thereby operationalising the data minimisation obligation from the outset.
The CNIL expressly rejects the argument that purpose limitation is fundamentally incompatible with AI model training, which may develop unanticipated emergent characteristics. The requirement to define a purpose must be adapted to the context of AI, but it does not disappear. The authority distinguishes three practical scenarios.
Three Practical Scenarios
First, where the developer clearly knows the operational use of the system, the operational objective will serve as the purpose for both the development and deployment phases. Second, for general purpose AI (GPAI) systems — for example, large language models or generative AI systems for images, video, or code — it is not permissible to define the purpose so broadly as to encompass merely “the development and improvement of an AI system.” The developer must instead specify the type of system being developed and its technically feasible functionalities and capabilities. Third, for AI systems developed for scientific research purposes, a less detailed objective may suffice at the outset, provided additional information is supplied as the project progresses.
Practical note: Counsel advising AI developers should prioritise the early preparation of a precise purpose statement. This document is the cornerstone of downstream compliance decisions, including the choice of legal basis, the scope of data minimisation obligations, and the boundaries of permissible data re-use.
Step 2 Determine Your Responsibilities
The Principle
Before undertaking any personal data processing in connection with AI development, an organisation must determine its role under the GDPR. The two primary roles are data controller (the entity that determines the purposes and means of processing) and data processor (an entity that processes data on behalf of, and strictly under the instructions of, a controller). Where two or more entities jointly determine purposes and means, they are joint controllers and must define their respective obligations, typically by contract.
The AI Act Overlay
The EU AI Act introduces its own role taxonomy: AI system providers (entities that develop or place AI systems on the market), importers, distributors, and deployers. These AI Act roles do not map neatly onto the GDPR controller/processor distinction, and the degree of GDPR responsibility requires a case-by-case analysis. The CNIL identifies the following illustrative scenarios:
- Provider as data controller: Where a provider initiates AI system development and constitutes the training dataset from self-selected data, it will generally qualify as a data controller.
- Joint controllers: Where a provider builds a training dataset together with other data controllers for a jointly defined purpose, the parties will be joint controllers.
- Provider as data processor: Where a provider develops an AI system on behalf of a customer who determines both the purpose and the technical means, the provider may be classified as a data processor. If the customer specifies only the goal and leaves system design to the provider, the provider retains data controller status.
- Subcontractors: A service provider engaged by an AI system provider to collect and process data according to precise instructions will be classified as a subcontractor (data processor) of that provider.
Data processors operating in an AI development context carry specific obligations: they must operate under a data processing agreement compliant with Article 28 GDPR, follow controller instructions strictly, maintain the security of processed data, and assess compliance at their own level, alerting the controller if problems arise.
Step 3 Define the Legal Basis
The Six Available Legal Bases
The development of AI systems involving personal data requires a valid legal basis under Article 6 of the GDPR. The six available bases are: consent; compliance with a legal obligation; performance of a contract; performance of a task in the public interest; safeguarding vital interests; and the pursuit of a legitimate interest. The choice of legal basis has significant downstream consequences for the nature of obligations owed to data subjects and the rights they may exercise, making early and careful selection essential.
Consent
Where data is collected directly from individuals who can freely accept or refuse without detriment, consent is often the most appropriate legal basis. To be valid, consent must be freely given, specific, informed, and unambiguous. In practice, however, consent is frequently unavailable — for example, when training data is sourced from publicly accessible online content or from pre-existing open source datasets, where direct contact with data subjects is not feasible.
Legitimate Interest
Legitimate interest is among the most commonly invoked legal bases for AI system development, particularly by private actors. Its use is subject to a three-part test. First, the interest pursued must be legitimate — lawful under all applicable legislation (including the AI Act), clearly and specifically defined, and linked to the organisation’s mission and activities. Examples of interests generally considered legitimate by the CNIL include conducting scientific research, facilitating public access to certain types of information, offering a conversational assistant service, and developing an AI system to detect fraudulent content or behaviour. A commercial interest may qualify as legitimate provided it is lawful and the processing is necessary and proportionate. Second, the processing must be necessary — the pursued interest cannot be achieved by less privacy-intrusive means, and the necessity of the processing must be assessed against the data minimisation principle. Third, the processing must not cause a disproportionate impact on individuals’ privacy, which requires the controller to weigh the expected benefits of the processing against its potential impact on the rights and freedoms of the individuals concerned.
Reasonable Expectations and Safeguards
The CNIL places significant emphasis on individuals’ reasonable expectations when relying on legitimate interest. The use of personal data should not come as a surprise. For data collected directly from individuals, relevant factors include the relationship with the individual, the context of collection, the nature of the service, and privacy settings. For the reuse of online data, the authority warns that processing will not fall within reasonable expectations where the scraping includes websites that have set restrictions via terms of use, robots.txt files, or CAPTCHA protections.
Where the disproportionate impact test is not satisfied without additional measures, the CNIL outlines a menu of safeguards capable of reducing the impact of processing to an acceptable level. These include: timely anonymisation or pseudonymisation of collected data; measures to reduce the risk of model memorisation and subsequent data extraction or regurgitation; a discretionary and prior right to object; the right to have personal data erased from training datasets; mechanisms for individuals to be identified when exercising their rights; and active communication regarding updates to datasets or models.
Public Actors
Public actors must verify that their processing is in line with a public interest mission as provided for by law (a statute, decree, or equivalent) and that it contributes to that mission in a relevant and appropriate way. The CNIL gives the example of the French Pôle d’expertise de la régulation numérique (PEReN), which is authorised on this basis to reuse publicly available data for regulatory technology experiments.
Step 3 (bis) Adapt Safeguards to Web Scraping
Web Scraping Is Not Per Se Prohibited
The CNIL makes clear that web scraping is not, in itself, prohibited under the GDPR. Private actors may rely on legitimate interest as the legal basis for scraping, provided appropriate safeguards are implemented. This section of the recommendations is of particular practical significance for AI developers building training datasets from publicly accessible online content.
Data Minimisation in Practice
The CNIL requires scrapers to operationalise the data minimisation principle in three specific ways. They must define in advance the categories of data that are relevant before commencing collection; avoid collecting more than necessary by filtering or excluding certain categories of websites (particularly those likely to contain sensitive data); and delete any irrelevant data collected inadvertently without delay.
Respecting Reasonable Expectations and Mandatory Exclusions
Developers must take into account the public accessibility of data, the nature of source websites (for example, social media platforms versus open data platforms), and the type of content published. Web scraping that disregards technical protections — such as CAPTCHAs or robots.txt files — falls outside individuals’ reasonable expectations and is incompatible with legitimate interest processing.
Additional Recommended Safeguards
Depending on the intended use of the AI system and the associated risks, the CNIL recommends implementing one or more additional safeguards, including: establishing a default exclusion list of websites containing particularly sensitive data (such as health-related forums); limiting data collection to freely accessible content not requiring account creation; informing individuals as widely as possible (for example, through online articles or social media posts); providing a discretionary and advance right to object before data collection begins, with a reasonable delay before model training commences; and anonymising or pseudonymising data immediately after collection while preventing re-identification through user identifiers.
Step 4 Check Whether You Can Re-use Certain Personal Data
The Compatibility Test
Where a developer plans to re-use a dataset that was originally collected for a different purpose, it must verify that such re-use is lawful. Unless the data subjects have consented to the new use or the re-use is authorised by law, the developer must conduct a “compatibility test”. This test must take into account: the existence of a link between the initial objective and the objective of building an AI training dataset; the context in which the personal data were originally collected; the type and nature of the data; the possible consequences for the individuals concerned; and the existence of appropriate safeguards such as pseudonymisation. Where re-use is for statistical or scientific research purposes, compatibility is presumed by law and no test is required.
Re-use of Publicly Available (Open Source) Data
Where a provider re-uses a publicly available dataset, the CNIL recommends documenting in the DPIA that: the dataset’s description identifies its source; the dataset is not manifestly the product of a crime or offence and has not been the subject of a public sanction imposing removal; there is no serious doubt as to the lawfulness of the dataset; and the dataset does not contain sensitive data (or, if it does, that additional checks have confirmed the lawfulness of the sensitive data processing). While the original publisher of the dataset bears primary responsibility for its lawfulness, re-users retain an independent obligation to conduct these verification checks.
Re-use of Third-Party Data (Data Brokers)
Where data is acquired from third parties such as data brokers, the applicable rules depend on whether the third party originally collected the data for the purpose of building an AI training dataset (in which case it bears the primary compliance obligation) or for a different purpose (in which case it must have already conducted a compatibility test before transferring the data). In either scenario, the re-user must verify that it is not re-using a manifestly unlawful dataset. A contractual agreement between the original data holder and the re-user is strongly recommended to facilitate these verifications.
Step 5 Minimise the Personal Data You Use
The Principle
The GDPR requires that personal data collected and used be adequate, relevant, and limited to what is necessary in light of the defined purpose. The CNIL emphasises that this principle of data minimisation must be applied with particular rigour when sensitive data is involved. The principle does not prohibit training on large volumes of data, but it does require upstream reflection to identify and collect only the personal data genuinely useful for development.
Practical Implementation
The CNIL’s practical guidance on minimisation covers four areas. On the method to be used, developers should favour techniques that achieve the desired result using as little personal data as possible; the use of deep learning should not be systematic where less data-intensive methods are available. On selection of strictly necessary data, the minimisation principle requires an upstream analysis to identify relevant data before collection and the implementation of technical means to collect only those data. On validation of design choices, the CNIL recommends conducting a pilot study (which may use fictitious, synthetic, or anonymised data) and consulting an ethics committee or ethical advisor. On the organisation of the collection process, recommended steps include data cleaning to ensure quality and consistency; identification of relevant data to optimise performance while avoiding under- and over-fitting; the application of data protection by design measures (including generalisation and anonymisation); ongoing monitoring and updating of datasets to prevent data drift; and documentation of the data used to ensure traceability throughout the system’s lifecycle.
Step 6 Set a Retention Period
The Principle
The storage limitation principle under the GDPR prohibits the indefinite retention of personal data. Developers must define a retention period calibrated to the purpose that justified the processing. This obligation applies distinctly to: (1) data retained during the development phase (which must be pre-planned, monitored, and communicated to data subjects); and (2) data retained after completion of the development phase for maintenance or improvement purposes (which should in principle be deleted unless specific guarantees are in place, such as partitioned access restricted to authorised personnel).
The CNIL acknowledges that the retention of training data may be justified where it enables bias measurement or audit procedures. In such cases, prolonged retention is permissible provided it is limited to the data strictly necessary for those purposes and accompanied by enhanced security measures. Where general statistical information about the dataset (such as documentation of its statistical distribution) would suffice, individual data should not be retained.
Step 7 Inform Individuals
The Information Obligation
The transparency principle under Articles 13 and 14 of the GDPR requires that individuals be informed about how their data will be used (the why, how, and manner of processing) so that they can exercise their rights, including the rights to object, access, and rectification. This obligation applies both to data collected directly from individuals and to data collected indirectly, including through web scraping.
Accessibility and Clarity
Information must be easily accessible: it may be provided at the individual level directly on a data collection form or via pre-recorded voice message, or at the general level on the developer’s website as part of a privacy notice. The CNIL particularly recommends allowing a reasonable delay between individual notification about data collection and the commencement of AI model training, given the difficulty individuals face in exercising their rights once a model has been trained. Information must be concise and clear, and the complexity of AI systems must not be used as a justification for unintelligible disclosures. The CNIL recommends illustrating, through diagrams or equivalent means, how data is used during training, how the AI system works, and how to distinguish between the training dataset, the AI model, and its outputs.
Exceptions and the Individual vs. General Information Choice
Notwithstanding the default obligation to provide individual-level information, the GDPR permits exceptions where the data subject already has the information, or where providing it would require disproportionate effort. The CNIL notes that individual notification is often considered disproportionate when pseudonymised data has been collected via web scraping, as identifying and contacting individuals may require collecting additional or more identifiable data. In such cases, a comprehensive general information notice published on the developer’s website is the recommended alternative.
The CNIL also draws attention to Article 53 of the AI Act, which requires providers of general-purpose AI models to complete a public summary of the content used to train the AI system, using a template provided by the AI Office. This summary may contribute to satisfying the general information obligation concerning data sources.
Specific Information Requirements for AI
In addition to the standard GDPR information requirements, AI-specific disclosures apply. Where data sources are limited, the precise identity of those sources must generally be disclosed. Where a large number of sources are involved (as in large-scale scraping exercises), it is acceptable to indicate the categories of sources with illustrative representative examples. Where a dataset or AI model is reused, the CNIL recommends providing a way to contact the original data controller, particularly where the dataset or model presents significant risks to data subjects. Where it is not possible to identify individuals within the training dataset or model, data subjects must be informed of this impossibility and guided as to what additional information they can provide to facilitate their identification if they wish to exercise their rights.
Step 8 Ensure the Exercise of Data Subject Rights
The Principle
Individuals retain the right to access, rectify, erase, restrict, and port their personal data in relation to both the training dataset and the trained AI model itself, provided the model is not considered anonymous. The challenges that AI systems pose to the exercise of these rights — including the difficulty of identifying a specific individual within a large dataset or a trained model — do not excuse the developer from implementing appropriate responses.
Right of Access
The right of access entitles any individual to obtain, free of charge, a copy of all personal data processed about them. Disclosure must not infringe on the rights and freedoms of others (including the rights of other data subjects, intellectual property rights, or trade secrets). Where data was obtained from a data broker, the right of access includes any available information regarding the source of the data.
Rights to Rectification, Erasure, and Objection
The exercise of data subject rights over an AI model is not absolute. The proportionality of the required measures depends on the sensitivity of the data, the risk of regurgitation or disclosure, and the impact on the organisation’s freedom of enterprise. The default mechanism for satisfying erasure or rectification requests affecting an AI model is retraining the model, which may take up to three months depending on complexity and volume. Where retraining is disproportionate, the CNIL recommends implementing encapsulation filters applied around the AI system, provided these filters are demonstrably effective and robust.
Exemptions
Developers may refuse data subject rights requests that are manifestly unfounded or excessive, or where the exercise of the relevant right is excluded by French or European law, or where processing is carried out for statistical, scientific research, or historical research purposes.
Step 9 Securing Your AI System
The Principle
Article 32 of the GDPR requires data controllers to implement appropriate technical and organisational security measures. For AI development, the CNIL recommends conducting a risk analysis with particular attention to three areas: software development security; the creation and management of training datasets; and system maintenance. The CNIL strongly recommends documenting security measures in a DPIA (see below).
Security Measures Specific to AI Development
The CNIL identifies three security goals, each accompanied by a non-exhaustive list of relevant measures. To ensure the confidentiality and integrity of training data, developers should: check the reliability, quality, and integrity of training data sources and their annotations throughout the entire lifecycle; log and manage dataset versions; use synthetic or dummy data for security testing and integration; encrypt backups and communications; control access to non-open source data; anonymise or pseudonymise data; segregate sensitive datasets; and prevent loss of control over data through organisational measures. To guarantee the performance and integrity of the AI system, developers should: incorporate data protection by design considerations from the outset; use verified development tools, libraries, and pre-trained models (with attention to the risk of backdoors); favour verified import and storage formats; use a controlled, reproducible, and easily deployable development environment; implement a continuous development and integration process; document the system design and functionality; and conduct internal or third-party security audits, including simulated attacks on the AI system. To anticipate how the system will work, developers should inform users of the system’s limitations, provide sufficient information for interpreting outputs, ensure the possibility of shutting down the system, and control AI outputs using filters, reinforcement learning from human feedback (RLHF), or digital watermarking techniques.
The CNIL notes that the most likely risks to AI systems today relate not to the model itself but to other system components — such as backups, interfaces, and communications — and that exploiting a software vulnerability to access training data may be easier for an attacker than conducting a membership inference attack. General information systems security measures must therefore be properly implemented alongside AI-specific measures.
Step 10 Assess the Status of an AI Model
The Principle
An AI model is a statistical representation of the characteristics of the dataset on which it was trained. Academic research has demonstrated that, in certain cases, this representation is detailed enough to permit the extraction of training data. Consistent with EDPB Opinion 28/2024, the CNIL takes the position that AI models trained on personal data must, in most cases, be considered subject to the GDPR. To determine whether a given model falls under the scope of the GDPR — and therefore whether it can be considered anonymous — developers must conduct an assessment of the model’s status, using means reasonably likely to be used in practice, including re-identification attack tests.
Indicators Triggering Re-identification Attack Testing
The CNIL identifies three categories of indicators that suggest re-identification attack tests should be conducted. Regarding the training data: indicators include high identifiability and precision of the data, heterogeneity, rarity, or duplication within the dataset. Regarding the model architecture: indicators include a high ratio between the number of parameters and the size of training data, a risk of overfitting, or a lack of confidentiality guarantees during training (such as the absence of differential privacy mechanisms). Regarding the model’s features and use cases: indicators include objectives of reproducing data similar to training data (as in generative AI), or prior successful re-identification attacks on similar models.
Bringing a Non-Anonymous Model Within GDPR Scope
Where a model cannot be demonstrated to be anonymous, the developer may seek to bring the overall AI system outside the scope of the GDPR by embedding the model within a system that implements robust measures to prevent data extraction. Measures that may contribute to this outcome include: making access to or retrieval of the model impossible; implementing access restrictions; limiting output precision or filtering model outputs; and applying the security measures described in Step 9. These measures must themselves be tested through systematic adversarial attacks on the system.
Step 11 Comply with GDPR Principles During the Annotation Phase
The Principle
Annotation involves assigning labels or tags to data items to serve as the ground truth from which a model learns to classify or distinguish information. Annotation activities are subject to the same GDPR obligations as other aspects of AI development.
Data Minimisation in Annotation
Annotations must be limited to what is necessary and relevant for model training and the intended functionality. This includes data indirectly related to the system’s functionality where its performance impact is demonstrated or reasonably plausible, as well as contextual information useful for performance measurement, error correction, or bias evaluation. Where technical constraints prevent full compliance with the minimisation principle, the developer must be able to demonstrate efforts to use the most relevant annotated dataset and to remove irrelevant annotations.
Accuracy Principle
Annotations must be accurate, objective, and, where possible, up to date. Inaccurate or biased annotations may be reproduced by the AI system, leading to degrading or discriminatory outputs — a risk with significant liability implications for developers and deployers.
Sensitive Data in Annotations
Annotation that includes sensitive data as defined by Article 9 GDPR — such as data revealing health information, racial or ethnic origin, political opinions, or religious beliefs — is as a rule prohibited. Exceptions may apply in certain research contexts, subject to specific safeguards including the use of objective and factual criteria for labelling (for example, describing skin colour in terms of pixel values rather than ethnic origin), strict limitation to information present in the data without interpretation, enhanced security for annotated data, and risk assessment for data regurgitation and inference. The CNIL underscores that the use of sensitive data should be avoided wherever possible and replaced with synthetic data where feasible.
III. Focus: The Data Protection Impact Assessment (DPIA)
When Is a DPIA Required?
The GDPR requires a DPIA under Article 35 where processing is likely to result in a high risk to the rights and freedoms of individuals. For AI development, the CNIL “strongly recommends” conducting a DPIA where at least two of the following criteria are met: sensitive data is collected; personal data is collected at large scale; data of vulnerable persons (minors, persons with disabilities) is collected; datasets are crossed or combined; or new technological solutions are implemented or innovative use is made. Even where fewer than two criteria are met, a DPIA is mandatory if significant risks exist, such as the risk of data misuse, data breach, or automated discrimination.
For AI Act high-risk AI systems involving personal data, the CNIL considers that a DPIA is in principle required under the GDPR. The DPIA may be based on the technical documentation already required by the AI Act, provided it incorporates all elements required under Article 35 GDPR.
Scope of the DPIA
The scope of the DPIA depends on whether the developer knows the operational use of the system at development time. Where operational use is known, a general DPIA covering the entire lifecycle (development and deployment) is recommended, with the deployer responsible for the deployment-phase DPIA (which may be based on the developer’s DPIA). Where a general purpose AI system is being developed, only a development-phase DPIA is feasible; this DPIA should be made available to users of the system to enable them to conduct their own deployment-phase analysis.
AI-Specific Risks to Address in the DPIA
The CNIL identifies a comprehensive set of AI-specific risks that must be assessed and addressed in the DPIA: risks to data confidentiality arising from possible extraction of personal data from the AI system; risks linked to misuse of training dataset content by employees with access; the risk of automated discrimination caused by biases introduced during development; the risk of generating false content relating to a real person (particularly in generative AI); the risk of automated decision-making in contexts where operators cannot verify system performance or override its outputs; the risk of users losing control over their publicly available personal data; risks from known AI-specific attacks (such as data poisoning); and systemic and serious ethical risks related to system deployment.
Mitigation Measures
Once the level of risk has been determined, the DPIA must specify a package of measures to reduce and maintain risk at an acceptable level. Available measures include: technical security measures (such as homomorphic encryption or secure execution environments); minimisation measures (such as the use of synthetic data); anonymisation or pseudonymisation measures (such as differential privacy); privacy by design measures from the outset (such as federated learning); measures facilitating the exercise of individual rights (such as machine unlearning techniques, explainability, and traceability mechanisms for system outputs); and audit and validation measures (such as simulated adversarial attacks). Organisational measures may also be implemented, including governance arrangements (such as the establishment of an ethics committee) and measures for the traceability of actions and internal documentation.
IV. Practical Takeaways
The CNIL’s recommendations constitute the most detailed and operationally comprehensive guidance yet issued by a major European supervisory authority on the intersection of the GDPR and AI system development. Several themes emerge as priority areas for AI developers and providers.
Purpose definition is foundational. Every downstream compliance decision — legal basis, data minimisation scope, retention period, information obligations, and data subject rights — flows from the purpose statement. Counsel should prioritise early engagement on purpose definition and ensure that it is documented, precise, and revisited as the project evolves.
The GDPR and the AI Act are complementary, not alternatives. Organisations that have already invested in AI Act compliance documentation (technical documentation, conformity assessments for high-risk systems) should map those deliverables to GDPR DPIA requirements to avoid duplicating effort and to ensure coherent compliance across both frameworks.
Legitimate interest requires rigorous three-part analysis. Organisations relying on this basis must document the legitimacy of the interest, the necessity of the processing, and the proportionality assessment, together with the safeguards implemented to mitigate disproportionate impact. Reliance on a cursory or boilerplate legitimate interest assessment will not suffice.
Web scraping is permissible but heavily conditioned. For developers building training datasets from publicly available online content, the CNIL’s Step 3(bis) guidance sets out a prescriptive regime. Counsel should advise clients to implement a documented scraping protocol that respects robots.txt and CAPTCHA restrictions, pre-defines data categories, excludes sensitive data sources, and provides for immediate deletion of inadvertently collected irrelevant data.
Model status assessment is not optional. The CNIL’s endorsement of EDPB Opinion 28/2024 on the GDPR applicability of AI models confirms that most AI models trained on personal data will remain subject to the GDPR even after training. Organisations should build re-identification attack testing into their development workflows and document the results as part of DPIA obligations.
Data subject rights over AI models require proactive infrastructure. The right of erasure, in particular, raises significant implementation challenges when applied to a trained model. Organisations should invest early in technical architectures that facilitate machine unlearning or model retraining, and should design contractual downstream obligations (through reuse licences and API terms) that propagate erasure and rectification obligations to downstream users and deployers.
Conclusion
The CNIL’s January 2026 recommendations represent a landmark contribution to the European data protection landscape for AI. They provide, for the first time in a single comprehensive document, a structured eleven-step compliance framework that takes AI development from initial purpose definition through to post-training model status assessment and annotation compliance. They also signal unambiguously that GDPR compliance is not an obstacle to AI innovation, but a framework within which responsible innovation can and must occur.
For practitioners, the recommendations are both a practical compliance guide and a harbinger of future enforcement. The CNIL and other European supervisory authorities have made clear that AI development involving personal data will be subject to rigorous scrutiny. Organisations that engage early and systematically with the framework outlined in these recommendations — and that document their compliance journey through robust DPIAs, purpose statements, and legitimate interest assessments — will be materially better positioned than those that do not.
