Sharing the Benefits of Data
AI can generate significant value from data while separating that value from the people, workers, communities, and public institutions that produce or steward it. Benefit sharing therefore requires attention not only to access to AI tools, but also to the terms under which data and labour enter AI value chains and how resulting benefits are distributed.
Data labelling / annotation and labour conditions: Labelling and annotation are routed through complex global supply chains that concentrate value at one end while dispersing risk and cost to the other, where workers are disproportionately located among the global majority. These workers collect, clean, classify and enrich the data necessary to train AI models while facing short-term and task-based contracts, low and inconsistent pay, intense algorithmic monitoring, and, in the case of content moderation, sustained exposure to distressing material. This labour should be considered in terms of benefit sharing, including through disclosure requirements that cover labour practices and provenance, minimum standards for data work, and value-distribution mechanisms that recognise the human contribution embedded in datasets before they enter the AI pipeline.
Data-to-AI value capture asymmetry: AI has become a major mechanism through which data is converted into economic value. Data generated by individuals, communities, and public institutions is aggregated and processed by a small group of actors with the technical and market capacity to train AI models, often without adequate consent, disclosure, or compensation mechanisms. This asymmetry means that the global majority, from whom the data often originates, typically does not receive their share of the value derived. Addressing the terms of data circulation through valuation methods, provenance disclosure, and licensing and benefit-sharing arrangements established before or when data enters AI pipelines is key to addressing this asymmetry.
Data trusts, cooperatives, collective governance institutions: Individual consent-based frameworks are inadequate to govern the cumulative value and harms created when data is aggregated for AI. Collective governance institutions, such as data trusts and cooperatives, can give communities collective authority to negotiate access, determine acceptable use, and influence how value is distributed.
Public sector data feeding AI development: Open data and other publicly available public-sector data is increasingly being accessed or licensed for use by AI developers.2 The governance of this access concerns control and value capture from data that has traditionally been stewarded by the State. A core consideration, therefore, is whether the terms governing this use reflect applicable consent requirements, value-sharing and licensing arrangements, as well as privacy, data protection, security and intellectual property rules. Privacy and data protection frameworks are particularly important to ensure open data does not disclose or enable the disclosure of personal information where not permitted by law. When terms of service do not fully account for applicable rules, or the public-domain status of data is unclear, governments may lack a clear and deliberate policy position on how open or publicly available public sector data should be treated per use case.
IP, licensing, attribution for data used in AI training: How intellectual property, licensing, and attribution rules apply to data used to train AI models determines whether the creators and rightsholders whose works constitute training data share in the value those models generate. Fragmented national rules and case-by-case copyright exceptions are shifting attention toward licensing, remuneration, and attribution mechanisms for training data. These risks can be mediated through provenance attribution, clarifying applicable rights and licenses, and supporting voluntary collective, or statutory compensation arrangements that benefit individual creators and source communities—not only to large rightsholders or intermediaries.
Data-related capacity, dependency and institutional power: States’ ability to negotiate the terms of data use in the AI pipeline depends on their legal, regulatory, and institutional capacity to steward data, scrutinize access agreements, and retain value. Global cooperation should support contestability: the ability of groups, communities, institutions, and States to identify and challenge potential violations of data-governance principles and seek appropriate remedy. This may require consideration of procedures for trans-jurisdictional claims, the standing of data creators and originators, and the circumstances in which rights over data may be asserted or its release refused. Such mechanisms can help address asymmetries that otherwise limit countries’ influence over how their data is used and how resulting value is distributed.
2. Global Partnership on Artificial Intelligence (GPAI), The Role of Government as a Provider of Data for Artificial Intelligence: Phase 1 Full Report (May 2024), https://wp.oecd.ai/app/uploads/2025/05/role-of-government-as-a-provider-of-data-for-AI-phase-1-full-report-1.pdf.