Skip to content

Glossary AI content labelling

What is AI content labelling

Definition

AI content labelling is the indication, visible to the public or readable by machine, that a text, an image, an audio file or a video has been created or modified with an artificial intelligence system.

On this page 5
  1. What labelling means and what is being labelled
  2. What the law requires and what Google requires
  3. Why it matters
  4. Good practice
  5. Common mistakes
In brief

Labelling AI content means indicating how a piece was made, through a note visible to the reader or a machine-readable mark inside the file.

What labelling means and what is being labelled

Three different things live under a single word. The first is the visible note addressed to the reader, along the lines of image generated with artificial intelligence. The second is the machine-readable marking that the model provider embeds in the output, meant for a later system to recognise it. The third is documented provenance, a signed history recording what captured or generated the file and which edits it has undergone since.

The techniques map onto those layers. The Content Credentials of the C2PA standard keep that signed history inside the file. Invisible watermarks alter pixels, waveform or word choice imperceptibly so that the signal travels in the content rather than in the metadata. Classic image metadata, in formats such as IPTC or XMP, offers a field describing the creation process.

How well each layer survives varies a great deal. Metadata and embedded manifests are easily lost when a file goes through a platform's processing, a screenshot or a recompression, which is why C2PA itself falls back on watermarks and fingerprints to reconnect the file with its credential. None of these techniques is a detector: they confirm what a system declared, not what a file carrying no signal actually is.

What the law requires and what Google requires

Regulation (EU) 2024/1689 on artificial intelligence concentrates transparency obligations in Article 50, applicable since 2 August 2026. Paragraph 2 requires providers of systems generating synthetic audio, image, video or text to mark the output in a machine-readable format and make it detectable as artificially generated or manipulated, as far as this is technically feasible.

Paragraph 4 addresses whoever deploys the system and requires disclosure of the artificial origin in two cases: deepfakes, and text published in order to inform the public on matters of public interest. The second case falls away where the content has undergone human review or editorial control and a natural or legal person holds editorial responsibility. Paragraph 5 asks that the information be provided in a clear and distinguishable manner, at the latest at the time of first exposure. Spain is also processing an artificial intelligence governance law with penalties for breaching labelling duties, whose status is worth checking before publishing.

Google is where the industry gets most confused. Its documentation does not ask for AI-generated content to be labelled and does not penalise a page for having been written with a model. What its spam policy sanctions is scaled content abuse, defined as producing many pages whose primary purpose is to manipulate rankings, regardless of how they were created. The helpful content guidance does suggest explaining how a piece was produced when the reader might wonder about it, and frames that as an editorial recommendation rather than a ranking requirement.

Why it matters

The decision that depends on this is twofold. First, whether a specific piece falls into a mandatory case, because that determines the form of the label, its placement and the moment it must appear. Second, how much the organisation invests in leaving a trace of the process, which costs money in workflows and training and which no tool solves on its own.

Separating the two layers avoids the most common error. When a team believes Google requires the label, it ends up putting a generic notice at the foot of every article that informs nobody and covers no legal case. And when a team believes the label harms rankings, it hides the process precisely in the pieces where the audience does expect to know, such as analyses, comparisons or news.

For a newsroom working in several languages and several markets there is a third effect. Documented human review is not only the exception provided for informative text, it is also the only way to catch invented claims before they go out. The label describes where a piece came from, while the review answers for its content.

Good practice

  • Sort your formats before drafting any policy: separate informative pieces on matters of public interest, pieces containing synthetic people or voices, and purely commercial ones, because the applicable regime is not the same.
  • Record the human review with a name, a date and the scope of what was reviewed, since it is the exception provided for informative text and only counts if it can be evidenced.
  • Write the note into the piece rather than into a general policy: the reader must find it on first contact with the content, not in a legal notice three clicks away.
  • Keep the original file with its provenance credentials in your own storage, because the copy published on an external platform usually loses them.
  • Distinguish the editorial note about the process from the legal obligation, and explain that difference in writing to the content team so that nobody decides by intuition.
  • Have the mandatory cases reviewed by legal counsel when you publish in several markets, because the classification depends on the type of content and on who holds editorial responsibility.

Common mistakes

  • Repeating that Google penalises AI content. Its policies address content created at scale to manipulate rankings, regardless of who or what wrote it.
  • Adding a generic notice at the foot of every page and treating that as compliance. A formula that identifies neither what was generated nor to what extent informs nobody.
  • Using an AI detector as evidence, on your own or someone else's text. These tools produce false positives and negatives, and no legal framework recognises them as a means of proof.
  • Assuming an image keeps its credentials after being published on a social network, when platform processing usually strips the metadata.
  • Confusing machine-readable marking with a visible label and delivering only one of the two when the case requires both.
Manuel Riveiro Rodriguez CEO & Digital Strategist

A technical audit covers this and everything else in one pass.

Request an audit

Frequently asked

Does Google penalise content generated with artificial intelligence?

Not for having been generated. Its documentation focuses on quality rather than on the production method. What it does sanction is scaled content abuse, meaning the publication of many pages whose primary purpose is to manipulate rankings, something that can happen with texts written by people or by models.

Am I required to label the articles on my blog?

It depends on the content. The duty to disclose artificial origin applies to text published to inform the public on matters of public interest, and falls away where there is human review with editorial responsibility assumed. A commercial article usually falls outside, although the specific classification is worth confirming with legal counsel.

Since when do these obligations apply?

The transparency obligations in Article 50 of the European artificial intelligence regulation have applied since 2 August 2026. They reach both whoever develops the system, who must mark the output in machine-readable form, and whoever deploys it, in the deepfake and informative text cases.

Does an AI detector work as evidence?

No. These services estimate probabilities from statistical features of the text and get it wrong in both directions, especially with translations and edited texts. Neither the law nor search engines recognise them as proof, so basing an accusation or a defence on them is risky.

Do Content Credentials survive publication on social networks?

Often they do not. The processing many platforms apply on upload removes metadata and embedded manifests, and a screenshot always loses them. That is why the standard falls back on watermarks and fingerprints that allow the credential to be recovered when the file arrives bare.

Sources

  1. Text of Article 50 of the European artificial intelligence regulation, with the marking and disclosure duties and its date of application.
  2. Google post on AI-generated content, stating that it values quality rather than the production method.
  3. Google spam policies, with the definition of scaled content abuse that is independent of how the content was created.
  4. Helpful content guidance, which suggests explaining how a piece was produced when the reader might wonder about it.
  5. C2PA frequently asked questions on Content Credentials, metadata loss and the use of watermarks and fingerprints.