Copyright and training data
Be able to reason about text and data mining, licences and the rights in generated material.
Prerequisites
- DLicences and open datarequired
Intuition
Three questions that often get conflated:
- May you train on protected material?
- Who owns what the model generates?
- What happens if the model reproduces its training data?
The answers differ between jurisdictions and are partly undecided in the courts. But the direction in the EU is clearer than in many other places.
Formal
1. Training on protected material — the EU. The DSM Directive (2019/790) articles 3 and 4 give an exception for text and data mining (TDM). Article 3 covers research organisations and cultural heritage institutions for scientific purposes. Article 4 covers everybody else, but the rights holder can reserve their rights — and for content online the reservation has to be machine-readable (in practice robots.txt or metadata). If the rights holder reserves, a licence is needed.
In Sweden this is implemented in the Copyright Act. The US has no equivalent exception and instead tests «fair use» case by case — several large disputes are ongoing.
2. Ownership of generated material. In the EU and Sweden, copyright requires human creation. Purely machine-generated material probably has no copyright protection. If the human creative contribution is large enough (selection, editing, composition) the result can be protected — but only in that part. The US Copyright Office has taken the same line.
The practical consequence: an image you generated can probably be used by anybody. That is rarely what the client thinks.
3. Reproduction of the training data. If the model outputs something that is in practice a copy of a protected work, that can be infringement, however it came about. Which is why you run memorisation checks: a nearest-neighbour search in the training data, and filtering the output against known works.
In a training pipeline all of this means three concrete requirements: record the licence per source, respect machine-readable reservations when collecting, and keep the provenance so that a source can be removed afterwards if the legal position changes. AI-grafen's source register with a licence, a fingerprint and re-verification every 30 days is built for exactly that.
Mastery means
- Reasons about text and data mining and its exceptions
- Assesses the rights in generated material
- Handles licences in a training pipeline
Sign in to do the exercises and build your mastery up.
Sources
- DSM-direktivet (EU) 2019/790, art. 3–4 — EU legal act
- PRV — copyright (in Swedish) — myndighetsmaterial
- arXiv — Extracting Training Data from Diffusion Models — arXiv (open access; licence per article)