---
authoritative: false
representation: annotated-page
publisher: AstroKube
methodology: https://ai-act.astrokube.com/about/
source_verified_on: '2026-09-15'
site_content_updated_on: '2026-09-15'
title: 'Article 10: Data and data governance | EU AI Act | AstroKube'
description: 'Article 10, EU AI Act: 1. High-risk AI systems which make use of techniques involving the training of AI models with data shall be developed on the basis…'
language: en
source: https://ai-act.astrokube.com/law/art-10/
---

1.  [Start](https://ai-act.astrokube.com/)
2.  [The law](https://ai-act.astrokube.com/law/)
3.  Article 10

Chapter III · Section 2 · Requirements for high-risk AI systems

# Article 10 — Data and data governance

⟲ Amended · Regulation (EU) 2026/1744

▼ Primary text, verbatim. Our annotations appear below, visibly separated.

[1\.](https://ai-act.astrokube.com/law/art-10/#p-1) High-risk AI systems which make use of techniques involving the training of AI models with data shall be developed on the basis of training, validation and testing data sets that meet the quality criteria referred to in paragraphs 2, 3 and 4 of this Article and in Article 4a(1) whenever such data sets are used.

[2\.](https://ai-act.astrokube.com/law/art-10/#p-2) Training, validation and testing data sets shall be subject to data governance and management practices appropriate for the intended purpose of the high-risk AI system. Those practices shall concern in particular:

(a) the relevant design choices;

(b) data collection processes and the origin of data, and in the case of personal data, the original purpose of the data collection;

(c) relevant data-preparation processing operations, such as annotation, labelling, cleaning, updating, enrichment and aggregation;

(d) the formulation of assumptions, in particular with respect to the information that the data are supposed to measure and represent;

(e) an assessment of the availability, quantity and suitability of the data sets that are needed;

(f) examination in view of possible biases that are likely to affect the health and safety of persons, have a negative impact on fundamental rights or lead to discrimination prohibited under Union law, especially where data outputs influence inputs for future operations;

(g) appropriate measures to detect, prevent and mitigate possible biases identified according to point (f);

(h) the identification of relevant data gaps or shortcomings that prevent compliance with this Regulation, and how those gaps and shortcomings can be addressed.

[3\.](https://ai-act.astrokube.com/law/art-10/#p-3) Training, validation and testing data sets shall be relevant, sufficiently representative, and to the best extent possible, free of errors and complete in view of the intended purpose. They shall have the appropriate statistical properties, including, where applicable, as regards the persons or groups of persons in relation to whom the high-risk AI system is intended to be used. Those characteristics of the data sets may be met at the level of individual data sets or at the level of a combination thereof.

[4\.](https://ai-act.astrokube.com/law/art-10/#p-4) Data sets shall take into account, to the extent required by the intended purpose, the characteristics or elements that are particular to the specific geographical, contextual, behavioural or functional setting within which the high-risk AI system is intended to be used.

− A passage here was deleted by the amendment.

[6\.](https://ai-act.astrokube.com/law/art-10/#p-6) For the development of high-risk AI systems not using techniques involving the training of AI models, paragraphs 2, 3 and 4 of this Article and Article 4a(1) shall apply only to the testing data sets.

This text is meant purely as a documentation tool and has no legal effect. The Union's institutions do not assume any liability for its contents. The authentic versions of the relevant acts, including their preambles, are those published in the Official Journal of the European Union and available in EUR-Lex.

**Amended**

— Regulation (EU) 2026/1744

Passages marked with the accent edge in the primary text were inserted or replaced by the amendment.

Show the text as adopted, before the amendment

The authentic 2024 text of this provision, shown for comparison. It no longer states the law.

1\. High-risk AI systems which make use of techniques involving the training of AI models with data shall be developed on the basis of training, validation and testing data sets that meet the quality criteria referred to in paragraphs 2 to 5 whenever such data sets are used.

2\. Training, validation and testing data sets shall be subject to data governance and management practices appropriate for the intended purpose of the high-risk AI system. Those practices shall concern in particular:

(a) the relevant design choices;

(b) data collection processes and the origin of data, and in the case of personal data, the original purpose of the data collection;

(c) relevant data-preparation processing operations, such as annotation, labelling, cleaning, updating, enrichment and aggregation;

(d) the formulation of assumptions, in particular with respect to the information that the data are supposed to measure and represent;

(e) an assessment of the availability, quantity and suitability of the data sets that are needed;

(f) examination in view of possible biases that are likely to affect the health and safety of persons, have a negative impact on fundamental rights or lead to discrimination prohibited under Union law, especially where data outputs influence inputs for future operations;

(g) appropriate measures to detect, prevent and mitigate possible biases identified according to point (f);

(h) the identification of relevant data gaps or shortcomings that prevent compliance with this Regulation, and how those gaps and shortcomings can be addressed.

3\. Training, validation and testing data sets shall be relevant, sufficiently representative, and to the best extent possible, free of errors and complete in view of the intended purpose. They shall have the appropriate statistical properties, including, where applicable, as regards the persons or groups of persons in relation to whom the high-risk AI system is intended to be used. Those characteristics of the data sets may be met at the level of individual data sets or at the level of a combination thereof.

4\. Data sets shall take into account, to the extent required by the intended purpose, the characteristics or elements that are particular to the specific geographical, contextual, behavioural or functional setting within which the high-risk AI system is intended to be used.

5\. To the extent that it is strictly necessary for the purpose of ensuring bias detection and correction in relation to the high-risk AI systems in accordance with paragraph (2), points (f) and (g) of this Article, the providers of such systems may exceptionally process special categories of personal data, subject to appropriate safeguards for the fundamental rights and freedoms of natural persons. In addition to the provisions set out in Regulations (EU) 2016/679 and (EU) 2018/1725 and Directive (EU) 2016/680, all the following conditions must be met in order for such processing to occur:

(a) the bias detection and correction cannot be effectively fulfilled by processing other data, including synthetic or anonymised data;

(b) the special categories of personal data are subject to technical limitations on the re-use of the personal data, and state-of-the-art security and privacy-preserving measures, including pseudonymisation;

(c) the special categories of personal data are subject to measures to ensure that the personal data processed are secured, protected, subject to suitable safeguards, including strict controls and documentation of the access, to avoid misuse and ensure that only authorised persons have access to those personal data with appropriate confidentiality obligations;

(d) the special categories of personal data are not to be transmitted, transferred or otherwise accessed by other parties;

(e) the special categories of personal data are deleted once the bias has been corrected or the personal data has reached the end of its retention period, whichever comes first;

(f) the records of processing activities pursuant to Regulations (EU) 2016/679 and (EU) 2018/1725 and Directive (EU) 2016/680 include the reasons why the processing of special categories of personal data was strictly necessary to detect and correct biases, and why that objective could not be achieved by processing other data.

6\. For the development of high-risk AI systems not using techniques involving the training of AI models, paragraphs 2 to 5 apply only to the testing data sets.

[

Recital 67 — interpretive context

High-quality data and access to high-quality data plays a vital role in providing structure and in ensuring the performance of many AI systems, especially when techniques involving the training of models are used, with a view to ensure that the high-risk AI system performs as intended and safely and it does not become a source of discrimination prohibited by Union law. High-quality data sets for training, validation…

](https://ai-act.astrokube.com/law/recital-67/)[

Recital 69 — interpretive context

The right to privacy and to protection of personal data must be guaranteed throughout the entire lifecycle of the AI system. In this regard, the principles of data minimisation and data protection by design and by default, as set out in Union data protection law, are applicable when personal data are processed. Measures taken by providers to ensure compliance with those principles may include not only anonymisation…

](https://ai-act.astrokube.com/law/recital-69/)

## What this means for you

In your terms · Data and data governance

Dataset lineage is yours: where each set came from, how it was prepared, what its gaps are, and the bias examination that was actually run.

-   Dataset cards with provenance
-   Bias examination report

Failure smells likeSomeone asks which data the model was trained on and the answer is a bucket path, with no record of how the set was assembled or what it was assumed to represent.

## Obligations derived from this article

[Data and data governanceArt. 10](https://ai-act.astrokube.com/explorer/?art=art-10)

## If you would rather not read the law

The basics page explains the Regulation's own categories in order: scope, role, tier, date. The engineering view groups the obligations by the platform capability they demand.

[Start with the basics →](https://ai-act.astrokube.com/basics/) [Open the engineering view →](https://ai-act.astrokube.com/engineering/)

## About this provision

### Status

Upcoming 2 Dec 2027

Moved from ~2 Aug 2026~

### Regime

High-risk

### Binds

Provider

### Type

Article · Chapter III · Section 2

### Amended by

[Regulation (EU) 2026/1744](https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:02024R1689-20260727)

### Recitals

[67](https://ai-act.astrokube.com/law/recital-67/) [69](https://ai-act.astrokube.com/law/recital-69/)

### Related

[Article 9](https://ai-act.astrokube.com/law/art-9/) [Article 11](https://ai-act.astrokube.com/law/art-11/)

### Cited capture

regulation-2024-1689/en-2026-08-18.html sha256 8f0b656302f9864c…

[Authentic text (EUR-Lex) →](http://data.europa.eu/eli/reg/2024/1689/oj) [This version (EUR-Lex) →](https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:02024R1689-20260727)

### Machine readable

[/law/art-10.md](https://ai-act.astrokube.com/law/art-10.md) [/api/law.json](https://ai-act.astrokube.com/api/law.json)

### Found an error?

[Write to us →](https://astrokube.com/contact)
