Introduction
Ahead of the next meeting of the UN Commission on Science and Technology for Development’s (CSTD) multi-stakeholder Working Group on Data Governance at All Levels as Relevant for Development, this brief outlines the areas of the Working Group’s mandate where AI-specific considerations should be considered for inclusion in outputs and provides a basis for dialogue among members and observers of the working group.
Data governance is foundational to, yet distinct from, governing AI. As stated in the Working Group’s Zero Draft - Progress Report: “AI systems depend on data quality, provenance, transparency, and accountability, thus underscoring the need for coordinated governance approaches across both spheres.”
Bringing AI to the discussion at the next WGDG
- Data governance is foundational to, yet distinct from, governing AI. Data governance issues that are relevant to AI must be acknowledged as such and addressed as part of data governance. Otherwise, the risk increases that AI tools will cause harm to the rights, safety, and interests of impacted people and communities.
- Many of the data governance principles identified by the Working Group are applicable throughout the AI lifecycle. Understanding where the scale, complexity, and persistence of AI systems can make failures more difficult to detect, correct, or reverse will help strengthen data governance in the context of AI.
Data governance applies across the AI lifecycle, including to training data, model parameters, AI-generated outputs, and data produced through users’ interactions with AI systems. Existing standards treat data governance as a prerequisite for responsible AI development and deployment, rather than as a downstream component of AI governance.1
Deploying AI without adequate data governance threatens equitable access, system reliability, and the rights and safety of affected people and communities. Risks include biased or inaccurate outputs, privacy and security breaches, unlawful data use, and limited accountability or redress. The following sections examine priority AI-related data governance considerations under the four areas of the Working Group’s mandate.
1. See: ISO/IEC 42001:2023, Information technology — Artificial intelligence — Management system. See particularly Annex A, A.7 Data for AI systems. https://www.iso.org/standard/81230.html; National Institute of Standards and Technology (NIST), Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1, 2023, pp. 20–31.
Fundamental Data Governance Principles
Many of the data governance principles identified by the Working Group are applicable throughout the AI lifecycle. However, the scale, complexity, and persistence of AI systems can make failures more difficult to detect, correct, or reverse. The following principles are therefore especially important in the context of AI, and this would be useful to address in the recommendations:
Data provenance and lineage: Records of data’s origin, collection, licensing conditions, and subsequent transformations are necessary to trace AI training data to their source. Such documentation enables fairness audits, supports communities to assert control over how their data is used, distinguishes data governance failures from model design failures, and enables detection (and removal, if necessary) of synthetic or low-quality training datasets.
Data quality and representativeness: The accuracy, completeness, consistency, timeliness, and fitness for purpose of training data shape AI model performance. Low quality data risks reproducing errors at scale while representativeness (the degree to which data reflects what it is meant to describe) gaps can lead to unequal AI model performance across groups. Data governance addresses these problems at their source, rather than attempting solutions at the model layer.
Data privacy, data protection, purpose limitations: Data incorporated into trained AI models can be difficult or impossible to remove. AI models, even when trained on de-identified or aggregated data, have been shown capable of re-identifying individuals. Cross-border data protection gaps may become AI training location incentives, with training activity migrating to where protection regimes are weakest, making global data-for-AI governance approaches all the more critical.
Synthetic data and provenance: Fully or partially AI-generated data (rather than data collected from observation, report, or measurement of real-world events) can degrade AI model performance when re-entering the AI training data pipeline over successive generations and may undermine representativeness without detection. Global initiatives should support provenance and transparent labeling of AI-generated or AI-modified data and apply data governance requirements to synthetic data throughout its lifecycle.
Interoperability
AI increases both the value of interoperable data systems and the risks of pursuing interoperability without appropriate governance. Interoperability across data systems may also enable international cooperation, for example, on research and development related to AI, while also helping lower barriers to trade and investment. International cooperation should therefore support appropriate technical and governance foundations while allowing regulatory autonomy and implementation to reflect national contexts, priorities, and capacities.
Interoperability of National Statistical and administrative data systems: Common technical standards and metadata can enable locally produced data to be securely linked, compared, and reused across national, regional, and international data ecosystems. However, the regulatory and institutional arrangements governing access, sharing, and reuse should reflect national contexts, capacities, and political priorities. Global cooperation should therefore support technical interoperability while preserving national regulatory autonomy, recognizing that uniform rules for data flows could deepen existing inequalities.
Sharing the Benefits of Data
AI can generate significant value from data while separating that value from the people, workers, communities, and public institutions that produce or steward it. Benefit sharing therefore requires attention not only to access to AI tools, but also to the terms under which data and labour enter AI value chains and how resulting benefits are distributed.
Data labelling / annotation and labour conditions: Labelling and annotation are routed through complex global supply chains that concentrate value at one end while dispersing risk and cost to the other, where workers are disproportionately located among the global majority. These workers collect, clean, classify and enrich the data necessary to train AI models while facing short-term and task-based contracts, low and inconsistent pay, intense algorithmic monitoring, and, in the case of content moderation, sustained exposure to distressing material. This labour should be considered in terms of benefit sharing, including through disclosure requirements that cover labour practices and provenance, minimum standards for data work, and value-distribution mechanisms that recognise the human contribution embedded in datasets before they enter the AI pipeline.
Data-to-AI value capture asymmetry: AI has become a major mechanism through which data is converted into economic value. Data generated by individuals, communities, and public institutions is aggregated and processed by a small group of actors with the technical and market capacity to train AI models, often without adequate consent, disclosure, or compensation mechanisms. This asymmetry means that the global majority, from whom the data often originates, typically does not receive their share of the value derived. Addressing the terms of data circulation through valuation methods, provenance disclosure, and licensing and benefit-sharing arrangements established before or when data enters AI pipelines is key to addressing this asymmetry.
Data trusts, cooperatives, collective governance institutions: Individual consent-based frameworks are inadequate to govern the cumulative value and harms created when data is aggregated for AI. Collective governance institutions, such as data trusts and cooperatives, can give communities collective authority to negotiate access, determine acceptable use, and influence how value is distributed.
Public sector data feeding AI development: Open data and other publicly available public-sector data is increasingly being accessed or licensed for use by AI developers.2 The governance of this access concerns control and value capture from data that has traditionally been stewarded by the State. A core consideration, therefore, is whether the terms governing this use reflect applicable consent requirements, value-sharing and licensing arrangements, as well as privacy, data protection, security and intellectual property rules. Privacy and data protection frameworks are particularly important to ensure open data does not disclose or enable the disclosure of personal information where not permitted by law. When terms of service do not fully account for applicable rules, or the public-domain status of data is unclear, governments may lack a clear and deliberate policy position on how open or publicly available public sector data should be treated per use case.
IP, licensing, attribution for data used in AI training: How intellectual property, licensing, and attribution rules apply to data used to train AI models determines whether the creators and rightsholders whose works constitute training data share in the value those models generate. Fragmented national rules and case-by-case copyright exceptions are shifting attention toward licensing, remuneration, and attribution mechanisms for training data. These risks can be mediated through provenance attribution, clarifying applicable rights and licenses, and supporting voluntary collective, or statutory compensation arrangements that benefit individual creators and source communities—not only to large rightsholders or intermediaries.
Data-related capacity, dependency and institutional power: States’ ability to negotiate the terms of data use in the AI pipeline depends on their legal, regulatory, and institutional capacity to steward data, scrutinize access agreements, and retain value. Global cooperation should support contestability: the ability of groups, communities, institutions, and States to identify and challenge potential violations of data-governance principles and seek appropriate remedy. This may require consideration of procedures for trans-jurisdictional claims, the standing of data creators and originators, and the circumstances in which rights over data may be asserted or its release refused. Such mechanisms can help address asymmetries that otherwise limit countries’ influence over how their data is used and how resulting value is distributed.
2. Global Partnership on Artificial Intelligence (GPAI), The Role of Government as a Provider of Data for Artificial Intelligence: Phase 1 Full Report (May 2024), https://wp.oecd.ai/app/uploads/2025/05/role-of-government-as-a-provider-of-data-for-AI-phase-1-full-report-1.pdf.
Safe, secure and trusted data flows, including cross-border flows
AI changes the nature and consequences of data flows by allowing data transferred for one purpose or at one stage of the AI lifecycle to be retained, transformed, or reused in others. Safe, secure, and trusted data flows therefore require governance arrangements that account for how different categories of data move through AI systems and the distinct risks associated with their transfer, retention, and reuse across borders.
Distinguishing between input and inference data in cross-border data flows: Data used to train or fine-tune AI models create different governance requirements than data processed during inference, such as prompts and user inputs. Training can encode or retain information within model parameters and propagate it across subsequent model versions and systems, making redress, withdrawal or correction challenging. Inference data can be controlled more directly through processing, retention and deletion requirements, with additional safeguards needed where they are retained or reused for further training. These distinctions are particularly important where data is stored, processed or used to train AI systems in another jurisdiction and may therefore be subject to different legal frameworks.
Implications for global data governance
- Data provenance and lineage: Establish common expectations for documenting data provenance and transformations across the AI lifecycle.
- Data quality and representativeness: Promote context-appropriate standards and assessment methods for data quality and representativeness.
- Data privacy, protection and purpose limitation: Uphold privacy, data-protection, and purpose-limitation requirements throughout the AI lifecycle and across borders.
- Synthetic data and provenance: Promote provenance and transparent labeling of AI-generated or AI-modified data and apply data governance requirements to synthetic data throughout its lifecycle.
- Interoperability of national statistical and administrative data systems: Support interoperable national data systems through adaptable standards, sustained capacity-building, and safeguards governing access and reuse.
- Data labelling, annotation, and labour conditions: Treat human labour as part of benefit sharing, including through disclosure requirements covering labour practices and provenance, minimum standards for data work, and value-distribution mechanisms that recognize and compensate the human contribution embedded in datasets before they enter the AI pipeline.
- Data-to-AI value capture asymmetry: Govern the terms of data circulation through valuation methods, provenance disclosure, and licensing and benefit-sharing arrangements established before or when data enters AI pipelines.
- Data trusts, cooperatives, and collective governance institutions: Recognize these institutions as mechanisms for operationalizing collective governance related to data and benefit sharing.
- Public sector data used in AI development: Support States in establishing policies to govern how public sector data, including open data, are shared and used before they enter AI pipelines.
- IP, licensing, and attribution for data used in AI training: Require provenance and attribution, clarify applicable rights and licenses, and support voluntary, collective, or statutory compensation arrangements that benefit individual creators and source communities—not only large rightsholders or intermediaries.
- Data-related capacity, dependency, and institutional power: Strengthen national data stewardship and negotiating capacity, and ensure that groups, communities, institutions, and States can contest the terms and consequences of data use, assert their interests, and seek remedy.
- Distinguishing between training and inference data in cross-border data flows: Distinguish between these categories of data and establish corresponding requirements for cross-border transfer, retention, reuse, correction, and deletion.
Conclusion
Data governance and AI governance are distinct but interdependent. Addressing the issues specific to data governance identified in this brief through the Working Group’s existing mandate can help ensure that emerging global approaches to AI rest on accountable, inclusive, and development-oriented data-governance foundations.