Skip to main content
Compliance and Policy Sandbox

Working out the rules of responsible data reuse, together

A standing series of open round tables and policy pilots where OpenAIRE members, policy makers, digital rights experts and the wider community tackle one question: how do we keep open science open in the age of AI, without surrendering provenance, rights or control?
The Problem

Open repositories are being scraped to breaking point

82%

of AI-bot activity is now training-related, up from about 72% a year earlier.
Source: Cloudflare Radar, full-year 2025

40%

of a repository's bandwidth a single deep-crawl bot (CCBot, GPTBot) can consume.
Source: OpenAIRE AI Bots & Repositories brief, 2026

2.5M+

sites already block AI-training crawlers, openness forced into retreat to survive.
Source: Cloudflare, 2026
The Paradox

To survive AI scraping, open repositories are closing the very doors they exist to keep open.

They were built to serve the open science community, so that anyone can find and reuse research. Yet AI crawlers hit them at machine scale, and to stay online many now add login walls, IP blocks and disabled bulk access, the very measures that work against theiw own mission. T

he deeper cause is a missing layer: AI agents have no machine-readable rights information, so they cannot tell what they may use, and simply take everything. OpenREL is built to close exactly that gap, so repositories can stay open because they are governed, not closed because they are exposed.

What is happening

Open content is under pressure from two sides

AI has changed the economics of open access. The community needs one shared, practical response, not every institution improvising alone.

Repositories under crawler load

AI training crawlers now make up the majority of bot traffic on open repositories. Under the strain, many are forced into login walls, IP blocks and disabled bulk endpoints, the very measures that shut down openness in order to survive.

Content reused without provenance

Open outputs are pulled into proprietaty models with no attribution, consent or traceability. Repositories carry the full cost of scraping while capturing none of the control or the benefit.

A fragmented legal landscape

The AI Act, Data Act, Data Governance Act, the DSM Directive and database rights all apply, but implementation is uneven, and most organisations lack the legal capacity to turn them into everyday practice.

On the table

What we will work on together

A rolling agenda, shaped by the people in the room. Where we will start:

A fair channel for AI data

Give AI developers a proper way to get the data, so they stop scraping repositories to get it.

Set clear reuse rules

Use database rights, licenses and OpenREL to state who can reuse what, and how.

Permissions machines can read

Turn "please do not train on this" and "reuse is fine" into signals software can follow.

Ready-made templates

Shared agreements and license clauses any member can pick up and use.

Make EU law usable

Turn the AI Act, Data Act and FAIR principles into steps people can actually follow.

Know the real cost

Measure how much traffic is bots rather than people, and what it costs.

A tool to put to work

Co-shape OpenREL with the research community

Most of this comes down to the one practical gap: a license is text written for people, and machines cannot read it. Research is no longer read only by humans, AI agents now discover and consume it, and without structured rules they simply proceed. OpenREL closes that gap, and the sandbox is where we put it to work.

A standard CC BY label tells you

✅ The work is free to use

✅ Attribution is required

A standard CC BY label tells you

❌ Can an AI model train on this?

❌ What is the GDPR consent status?

❌ Are there embargo conditions?

❌ Which reuse contexts are allowed?

What is OpenREL?

A policy layer for research data, built by OpenAIRE and DANS in the EOSC Beyond project. It turns licences, access rules and ethical or legal conditions into structured, machine-readable policies, expressed three ways at once: human-readable, plain legal text, and machine-actionable JSON-LD. It is built on the W3C ODRL standard. It does not replace licenses, it makes them work in automated, AI-driven environments, returning a clear decision with a reason every time. (OpenREL by P. Tsiavos and M. Katsamakis (OpenAIRE), EOSC Beyond project, with support from the EOSC EU Node. Introduced at FOR2026, Munich, May 2026. Slides: doi.org/10.5281/zenodo.20083991 (CC BY 4.0).

Who should take part

Bring your problem to the table

The sandbox only works if the people living the problem are in the room.

Repository and data managers

The people absorbing crawler load today, and the ones who need defences that do not break legitimate access.

Legal, policy and licensing experts

To translate EU frameworks into workable clauses, and to keep the group on the right side of the law.

Funders and policymakers

To connect what is learned here to standards, mandates and the wider European policy conversation.

Industry and AI developers

The demand side. A sanctioned channel only works if the people who want the data help design it. 

NOADs and the community

National desks and open science advocates who carry the outcomes back to institutions across Europe.

Timeline

Community driven efforts

The AI Bots and Repositories Working Group runs until December 2026, and the OpenREL tool is due to reach its finalised version around the same time. Community participation is the link between them: members who test OpenREL give it the real-world feedback it needs to finalise, and in return get machine-readable rights they can put straight to work. This leads to a shared outcome: a policy document the community owns.

  • Q3 2026, Sandbox opens
  • Autumn 2026, Build and feedback
  • December 2026, Convergence
  • Outcome 2026, A community policy document
  • January 2027, A potential hackathon

Participate

Members and the wider community are warmly invited. Tell us the data-reuse or compliance challenge you are facing, and will put it on the agenda for the next round table.