Introduction

Ahead of the next meeting of the UN Commission on Science and Technology for Development’s (CSTD) multi-stakeholder Working Group on Data Governance at All Levels as Relevant for Development, this brief outlines the areas of the Working Group’s mandate where AI-specific considerations should be considered for inclusion in outputs and provides a basis for dialogue among members and observers of the working group. 

Data governance is foundational to, yet distinct from, governing AI. As stated in the Working Group’s Zero Draft - Progress Report: “AI systems depend on data quality, provenance, transparency, and accountability, thus underscoring the need for coordinated governance approaches across both spheres.”

Bringing AI to the discussion at the next WGDG

  1. Data governance is foundational to, yet distinct from, governing AI. Data governance issues that are relevant to AI must be acknowledged as such and addressed as part of data governance. Otherwise, the risk increases that AI tools will cause harm to the rights, safety, and interests of impacted people and communities.  
  2. Many of the data governance principles identified by the Working Group are applicable throughout the AI lifecycle. Understanding where the scale, complexity, and persistence of AI systems can make failures more difficult to detect, correct, or reverse will help strengthen data governance in the context of AI.

Data governance applies across the AI lifecycle, including to training data, model parameters, AI-generated outputs, and data produced through users’ interactions with AI systems. Existing standards treat data governance as a prerequisite for responsible AI development and deployment, rather than as a downstream component of AI governance.1

Deploying AI without adequate data governance threatens equitable access, system reliability, and the rights and safety of affected people and communities. Risks include biased or inaccurate outputs, privacy and security breaches, unlawful data use, and limited accountability or redress. The following sections examine priority AI-related data governance considerations under the four areas of the Working Group’s mandate.



1. See: ISO/IEC 42001:2023, Information technology — Artificial intelligence — Management system. See particularly Annex A, A.7 Data for AI systems. https://www.iso.org/standard/81230.html; National Institute of Standards and Technology (NIST), Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1, 2023, pp. 20–31.

Implications for global data governance

  • Data provenance and lineage: Establish common expectations for documenting data provenance and transformations across the AI lifecycle.
  • Data quality and representativeness: Promote context-appropriate standards and assessment methods for data quality and representativeness.
  • Data privacy, protection and purpose limitation: Uphold privacy, data-protection, and purpose-limitation requirements throughout the AI lifecycle and across borders.
  • Synthetic data and provenance: Promote provenance and transparent labeling of AI-generated or AI-modified data and apply data governance requirements to synthetic data throughout its lifecycle.
  • Interoperability of national statistical and administrative data systems: Support interoperable national data systems through adaptable standards, sustained capacity-building, and safeguards governing access and reuse.
  • Data labelling, annotation, and labour conditions: Treat human labour as part of benefit sharing, including through disclosure requirements covering labour practices and provenance, minimum standards for data work, and value-distribution mechanisms that recognize and compensate the human contribution embedded in datasets before they enter the AI pipeline.
  • Data-to-AI value capture asymmetry: Govern the terms of data circulation through valuation methods, provenance disclosure, and licensing and benefit-sharing arrangements established before or when data enters AI pipelines.
  • Data trusts, cooperatives, and collective governance institutions: Recognize these institutions as mechanisms for operationalizing collective governance related to data and benefit sharing. 
  • Public sector data used in AI development: Support States in establishing policies to govern how public sector data, including open data, are shared and used before they enter AI pipelines.
  • IP, licensing, and attribution for data used in AI training: Require provenance and attribution, clarify applicable rights and licenses, and support voluntary, collective, or statutory compensation arrangements that benefit individual creators and source communities—not only large rightsholders or intermediaries. 
  • Data-related capacity, dependency, and institutional power: Strengthen national data stewardship and negotiating capacity, and ensure that groups, communities, institutions, and States can contest the terms and consequences of data use, assert their interests, and seek remedy.
  • Distinguishing between training and inference data in cross-border data flows: Distinguish between these categories of data and establish corresponding requirements for cross-border transfer, retention, reuse, correction, and deletion.

Conclusion

Data governance and AI governance are distinct but interdependent. Addressing the issues specific to data governance identified in this brief through the Working Group’s existing mandate can help ensure that emerging global approaches to AI rest on accountable, inclusive, and development-oriented data-governance foundations.