Content provenance (C2PA and watermarking)

What is content provenance?

Content provenance is the record of where a piece of media came from and what happened to it after that: which camera or which model made it, which software edited it, and who signs for that story. The record travels with the file, or can be looked up from the file, so that software can show it to whoever opens the image months later.

Most of this work happens under the C2PA standard, from the Coalition for Content Provenance and Authenticity. C2PA is a Linux Foundation project with more than 300 member organisations behind it, and the thing it defines is called a Content Credential: in the words of the specification, a cryptographically bound structure that records an asset's provenance. Sign a photo at the moment of capture and any later change to the pixels fails the check, so tampering shows up.

Be clear about what that buys you. C2PA's own explainer says provenance information alone cannot tell you whether the digital content is true, accurate or factual, and calls the standard no cure-all for misinformation. It tells you who claims to have made a file and whether that claim has been touched since. Whether the photographer staged the scene is a different question entirely.

The three mechanisms

Three techniques do the marking, and they fail in different places, which is why serious implementations combine them.

  1. Signed metadata. The generating or editing tool writes a manifest into the file: what made it, what was done to it, a hash of the actual pixels, and a signature from a certificate. A viewer recomputes the hash and checks the signature. This is what C2PA calls a hard binding, and it is exact. One changed pixel and the check fails.

  2. Invisible watermarking. The mark sits in the content itself rather than beside it. Google's SynthID nudges pixel values in images and video, adds an inaudible pattern to audio from Lyria and NotebookLM, and for text adjusts the probability scores of each generated word. Google says the image and video marks are built to survive cropping, filters, frame rate changes and lossy compression, and the audio marks to survive added noise, MP3 compression and speed changes. Reading the mark takes a detector that knows the algorithm.

  3. Fingerprinting. No mark is added at all. Software computes a perceptual hash of the content and looks it up in a repository that holds the provenance record. Because perceptual hashes match near-duplicates rather than exact copies, C2PA recommends showing anything found this way to a person for review instead of treating it as proof.

C2PA treats the second and third as soft bindings, and defines a Durable Content Credential as one that has a soft binding letting you find it again in a manifest repository. The reason is stated plainly in the specification: the credential may be routinely removed or corrupted during distribution, which happens on social platforms that serve resized or recompressed versions of an image without carrying the manifest along.

Signed metadata versus an invisible watermark

Take one dimension that decides most real cases: what happens when somebody screenshots your image and posts the screenshot.

The signed metadata is gone. A screenshot is a new file made by the operating system from the pixels on screen, and nothing carries over. The same thing happens on a smaller scale when a file is re-encoded, converted to another format or run through a tool that does not know about C2PA. OpenAI puts it in one line in its developer documentation: editing, converting or sharing a file can remove its metadata.

The invisible watermark often survives, because it lives in the pixels the screenshot copies. That is exactly why it exists. It is not guaranteed, since heavy cropping, filters and re-generation can wear it away, and it only helps if you have a detector for that specific watermark.

So the two are not competing options. The metadata carries the detail: which camera, which edits, which certificate. The watermark carries almost nothing, usually just an identifier, but it survives the trip. A file with both, plus a fingerprint in a repository, is the only combination with a real chance of still being recognisable after a few hops through a chat app.

Who has adopted this, and what you can check today

Cameras came first. The Leica M11-P, announced in October 2023, was the first camera to attach Content Credentials at capture. In September 2025 Google put the same thing in the Pixel 10 camera app, which reached C2PA Assurance Level 2, and Google Photos shows the credential and adds one when you edit.

On the software side, Adobe opened its free Content Authenticity web app in public beta in April 2025. You can sign up to fifty JPG or PNG files at once, whether or not they were made in an Adobe tool, attach a name verified through LinkedIn, and inspect other people's files with the Inspect tool or the Chrome extension. The Content Authenticity Initiative behind it passed 5,000 members in August 2025.

AI vendors mark their output both ways. OpenAI attaches C2PA Content Credentials to generated images and a SynthID watermark to supported images and audio, and offers a Verify check that looks for both in an uploaded file. Google reported in May 2026 that SynthID had marked over 100 billion images and videos and 60,000 years of audio, that OpenAI, Kakao, ElevenLabs and NVIDIA were adopting it, and that SynthID checking was live in the Gemini app and in Search, with Content Credentials checking following.

Platforms are the weak link, and the picture is uneven. TikTok was the first video platform to implement Content Credentials, in May 2024, reading them on upload to label AI content automatically and attaching its own. Meta announced in February 2024 that it would read the C2PA and IPTC markers on uploads to apply its AI label. Plenty of other places still strip everything on upload.

For a file in your hands right now, the practical route is the Verify page at contentcredentials.org, Adobe's Inspect tool, or the vendor's own checker if you suspect a specific tool made it.

What to watch out for with content provenance

  • Absence of a mark proves nothing. A negative result can mean the metadata was stripped, the file was re-encoded, the model predates provenance signals, or the content was made by a tool that marks nothing at all. OpenAI says outright that its checker is not a general-purpose AI detector. Read a clean result as no evidence found, never as evidence of no AI.

  • Removal is easy and undefended. The C2PA security document states that the standard offers no protection against the complete removal of manifests from assets. NIST's 2024 overview of synthetic content says the same about the whole category: covert and overt watermarks can often be removed, and embedded metadata can be stripped. That stripping is often not malicious. NIST notes that many platforms remove at least some metadata from uploads to protect privacy, and when provenance and privacy pull in opposite directions the platform usually picks privacy.

  • Detectors are probabilistic. Any covert watermark detector has a non-zero rate of false positives and false negatives. Treat a detection as one input to a judgement, not as the judgement.

  • A signature says who, not whether it is true. A correctly signed photograph of a staged scene is still a staged scene. Provenance moves the question from is this real to do I trust the entity that signed it.

Where the EU rules fit in

The technical layer has a legal counterpart. Article 50 of the AI Act has applied since 2 August 2026 and puts two separate duties on two separate parties: whoever provides a generative AI system has to mark its output so software can detect it, and whoever publishes a deepfake or AI-written public-interest text has to label it visibly. The Commission's Code of Practice on Transparency of AI-generated Content, finalised on 10 June 2026 and signed by roughly 190 organisations by the end of that July, describes how to do the first part and comes with a free set of EU icons for the second. The transparency obligation entry in this dictionary works through what that means for a company using these tools.

What to set up in a small company

Switch the feature on where you already have it. If you shoot on a Pixel or a recent Leica, or edit in Adobe tools, Content Credentials are a setting rather than a project. It costs you nothing and it gives a customer or a journalist a way to check your photo later.

Keep the originals. The signed file straight out of the camera is the version with the full record. Once it has been through a resize, a chat app and a website, it is a copy with no history. Archive the originals of your own photography and product shots, because that archive is what you fall back on if an image of yours is ever altered and republished.

Decide a house rule for your own AI images. Write down which categories of visual you label and how: a stylised illustration on a landing page is a different case from a photo-realistic image of a person or a place. Pick the rule once, put it in your style guide, and make it match what Article 50 asks of you rather than deciding image by image.

Ask suppliers what they mark. When you take on a generative AI tool, ask whether its output carries a machine-readable mark, which kind, and whether the vendor signed the Code of Practice. Note the answer next to the tool in your AI register, because when a question comes it will be about one specific file from one specific tool.

Last Updated: September 3, 2026 Back to Dictionary
Keywords
content provenance c2pa content credentials watermarking synthid ai transparency transparency obligation ai act metadata generative ai deepfake provenance