FROM AGPEDIA — AGENCY THROUGH KNOWLEDGE

AI content provenance

AI content provenance is information about where a piece of digital content came from and how it was made or changed, used in particular to show whether generative artificial intelligence produced or altered it. The main techniques are signed provenance metadata attached to a file, digital watermarks embedded in the content itself, and synthetic content detection tools that estimate after the fact whether content is machine-made. The US National Institute of Standards and Technology (NIST) groups these approaches under the term digital content transparency, which it defines as "the process of documenting and accessing information about the origins and history of digital content."[1:1]

Prominent provenance systems for AI output include the Coalition for Content Provenance and Authenticity (C2PA) metadata standard and watermarking systems such as Google DeepMind's SynthID.[2][3] From 2025, laws in China, the European Union, and California began to require that AI-generated content carry machine-readable markings.[4][5][6] Researchers have shown that current watermarks and metadata can be removed or forged, and NIST cautions that provenance information "may contribute to trustworthiness but does not guarantee it."[7][8][1:2]

Terminology and scope

"Provenance" also refers to the history of AI systems themselves: where a model's training data came from and how the model was built. That subject, sometimes called model provenance or data provenance, is distinct from the provenance of the content a model produces, which is the subject of this article.

Content provenance also covers more than AI. The same metadata standards can record that a photo was taken by a particular camera and edited in specific ways, which lets provenance assert that content is authentic as well as flag it as synthetic. NIST notes that provenance metadata can work "either by indicating synthetic origins or by asserting authenticity."[1:3] NIST uses the term "synthetic content" for "information, such as images, videos, audio clips, and text, that has been significantly altered or generated by algorithms, including by AI," a definition taken from US Executive Order 14110.[1:4]

Techniques

Technique How it works Main weaknesses
Provenance metadata Records origin and edit history alongside the file, often with a cryptographic signature Can be stripped or rewritten when files are copied or uploaded; hard to attach to plain text
Watermarking Embeds a hidden, machine-detectable signal in the content itself Can be weakened or removed by editing, regeneration, or paraphrasing; may be forged
Post hoc detection Classifier estimates whether content looks machine-generated Errors in both directions; false positives can harm people wrongly accused

Provenance metadata

Provenance metadata describes a piece of content's origin and history. It is usually stored inside the file it describes, though it can also sit in an external database linked to the content by an identifier.[1:3] The most widely discussed standard is the C2PA specification. The C2PA was formed on February 22, 2021, by Adobe, Arm, the BBC, Intel, Microsoft, and Truepic. It brought together the Adobe-led Content Authenticity Initiative and Project Origin, which Microsoft and the BBC led.[2] Version 1.0 of its specification, released on January 26, 2022, set out how to record information about images, video, audio, and documents so that "evidence of tampering can be identified."[9] As MIT Technology Review reporter Tate Ryan-Mosley explained, C2PA uses cryptography to record where content came from and to protect that record from tampering, using hashes that bind the provenance information to the content.[10]

Metadata travels poorly. NIST notes that metadata stored in a file "can similarly be stripped altogether, as it often is when files are shared (e.g., via social media platforms)," and that it is "not generally possible to have metadata travel with raw text as the text is copied across documents or applications."[1:5]

Watermarking

A digital watermark is a signal embedded in the content itself rather than stored beside it. Overt watermarks, such as a visible logo, are meant to be seen; covert watermarks are designed to be imperceptible to people but detectable by software. NIST observes that "unlike metadata, which is often stripped when content is disseminated, covert watermarks are designed to be persistent."[1:6]

For images, Google DeepMind launched SynthID in August 2023, initially for users of its Imagen image generator on Google Cloud. SynthID changes pixels in a way the eye cannot see, and the mark was designed to remain detectable after screenshots or edits such as resizing or rotation. Pushmeet Kohli, vice president of research at Google DeepMind, said it was more resistant to tampering than earlier image watermarks, though not perfectly immune. The launch followed voluntary commitments, announced by the White House in July 2023, in which companies including OpenAI, Google, and Meta agreed to develop watermarking tools.[3]

For text generated by large language models, watermarks work by nudging which words the model chooses. In a 2023 scheme by John Kirchenbauer and colleagues, the model selects a random set of "green" tokens (word pieces) before generating each word and slightly favors them; a detector then runs a statistical test for an unusually high share of green tokens. The detector does not need access to the model itself.[11] In 2024, Google DeepMind described SynthID-Text in Nature. It changes only the sampling step of generation, and the company reported a live experiment across nearly 20 million responses from its Gemini chatbot in which user feedback indicated that text quality was preserved.[12]

Post hoc detection

Detection tools try to classify content as human- or machine-made without relying on any mark placed at creation. They generally output a score that can be read as a probability. NIST notes that false positives, in which human work is judged to be AI-generated, "can be extremely damaging, potentially resulting in major reputational harms or adverse treatment."[1:7]

Regulation

Jurisdiction Measure Main provenance requirement In force
China Measures for Labeling AI-Generated Synthetic Content, with national standard GB 45438-2025 Visible ("explicit") labels plus metadata ("implicit") labels naming the provider September 1, 2025
European Union AI Act, Article 50 Machine-readable marking of synthetic audio, image, video, and text; disclosure of deepfakes August 2, 2026
California California AI Transparency Act (SB 942, amended by AB 853) Large providers must embed hidden disclosures in AI images, video, and audio and offer a free detection tool August 2, 2026

China

The Cyberspace Administration of China released final labeling measures in March 2025, effective September 1, 2025. They require visible labels on AI-generated content and "implicit" labels, defined as metadata containing details such as the service provider's name and a content ID. Content distribution platforms must check for these labels and mark content as confirmed, possible, or suspected AI-generated.[4]

European Union

Article 50(2) of the EU AI Act requires providers of AI systems that generate synthetic audio, image, video, or text to "ensure that the outputs of the AI system are marked in a machine-readable format and detectable as artificially generated or manipulated." Article 50(4) separately requires those who deploy systems to create deepfakes to disclose that the content is artificial.[5] The obligations apply from August 2, 2026.[5][13] A voluntary Code of Practice on transparency of AI-generated content, drafted by independent experts in a process run by the EU AI Office, was finalized on June 10, 2026.[13] According to a summary by the law firm Paul, Weiss, the code asks providers to mark outputs with "at least two machine-readable techniques," and a later legislative package, the Digital Omnibus, pushed the Article 50(2) marking deadline for systems already on the market to December 2, 2026.[14]

California

California's SB 942 applies to providers of generative AI systems with more than one million monthly users. It requires them to offer a free AI detection tool and to include hidden ("latent") disclosures in AI-generated image, video, and audio content; it does not cover text. AB 853, signed in October 2025, moved the start date from January 1, 2026, to August 2, 2026, and added duties for large online platforms from January 1, 2027, and for makers of cameras and other capture devices from January 1, 2028.[6]

Limitations

Research has shown that each technique can be defeated. Xuandong Zhao and colleagues showed in 2024 that invisible image watermarks that work at the pixel level can be removed by adding noise to an image and then reconstructing it with a generative model, while keeping image quality high.[7] For text, Vinu Sankar Sadasivan and colleagues found in 2023 that repeatedly paraphrasing AI-written text evaded watermark-based detectors as well as other kinds of detectors. They also showed that attackers could learn a watermark's pattern and "spoof" it, making human-written text appear machine-generated.[8]

Provenance also depends on adoption. Metadata only helps if the tools that create content add it and the platforms that display content preserve and show it; Ryan-Mosley noted that the C2PA standard is not legally binding and that "provenance labels do not necessarily mention whether the content is true or accurate."[10] People who want to deceive can use AI systems that add no marks at all. Discussing AI-generated abuse imagery, NIST notes that malicious actors often use freely available models "from which they can easily remove safeguards," and that they could attach watermarks or metadata to real abuse imagery so that it appears AI-generated and is less likely to be investigated.[1:8]

NIST warns that transparency measures can "create a false sense of trust," for example when a genuine, correctly labeled image is shown out of context.[1:2] It also raises privacy concerns: covert watermarks that carry data could reveal information about the person who made the content without that person's knowledge.[1:6]

Analysis

Provenance tools can support human agency by giving people more information about content they are deciding whether to trust, share, or act on. That benefit depends on how the information reaches people. Because metadata is easily stripped and watermarks can be removed or forged, a missing label is weak evidence that content is human-made, and a present label is not proof that content is accurate. If readers treat labels as verdicts rather than as one piece of evidence, provenance systems could narrow judgment instead of informing it, which is the "false sense of trust" NIST describes.[1:2]

Several questions remain unsettled and could be tested: whether watermarks can be made robust against regeneration and paraphrasing attacks without degrading quality, whether platforms will preserve provenance data once laws such as California's platform rules take effect in 2027, and how people actually interpret provenance labels in practice.

  1. ^ ↗ definition ^a ^b ^c ↗ trust-caveat ^a ^b ↗ metadata ^ ↗ synthetic-content ^ ↗ metadata-stripping ^a ^b ↗ covert-watermarks ^ ↗ false-positives ^ ↗ malicious-actors Chandra, Bilva; Dunietz, Jesse; Roberts, Kathleen; Lee, Yooyoung; et al. (2024-11-20). Reducing Risks Posed by Synthetic Content: An Overview of Technical Approaches to Digital Content Transparency. National Institute of Standards and Technology, Gaithersburg, MD. https://doi.org/10.6028/NIST.AI.100-4 https://www.nist.gov/publications/reducing-risks-posed-synthetic-content-overview-technical-approaches-digital-content.
  2. ^a ^b Microsoft (2021-02-22). Technology and media entities join forces to create standards group aimed at building trust in online content. Microsoft Source. https://news.microsoft.com/source/2021/02/22/technology-and-media-entities-join-forces-to-create-standards-group-aimed-at-building-trust-in-online-content/.
  3. ^a ^b Heikkilä, Melissa (2023-08-29). Google DeepMind has launched a watermarking tool for AI-generated images. MIT Technology Review. https://www.technologyreview.com/2023/08/29/1078620/google-deepmind-has-launched-a-watermarking-tool-for-ai-generated-images/.
  4. ^a ^b Luo, Yan (2025-03-18). China Releases New Labeling Requirements for AI-Generated Content. Inside Privacy (Covington & Burling). https://www.insideprivacy.com/international/china/china-releases-new-labeling-requirements-for-ai-generated-content/.
  5. ^a ^b ^c European Parliament and Council of the European Union (2024-06-13). Regulation (EU) 2024/1689 (Artificial Intelligence Act), Article 50: Transparency Obligations for Providers and Deployers of Certain AI Systems. https://artificialintelligenceact.eu/article/50/.
  6. ^a ^b Kye, Esther (2025-10-28). California AI Transparency Act Amendments Signed Into Law. Troutman Privacy + Cyber + AI. https://www.troutmanprivacy.com/2025/10/california-ai-transparency-act-amendments-signed-into-law/.
  7. ^a ^b Zhao, Xuandong; Zhang, Kexun; Su, Zihao; Vasan, Saastha; et al. (2024). Invisible Image Watermarks Are Provably Removable Using Generative AI. Advances in Neural Information Processing Systems (NeurIPS 2024). https://doi.org/10.48550/arXiv.2306.01953 https://arxiv.org/abs/2306.01953.
  8. ^a ^b Sadasivan, Vinu Sankar; Kumar, Aounon; Balasubramanian, Sriram; Wang, Wenxiao; et al. (2023). Can AI-Generated Text be Reliably Detected? Transactions on Machine Learning Research. https://doi.org/10.48550/arXiv.2303.11156 https://arxiv.org/abs/2303.11156.
  9. ^ Coalition for Content Provenance and Authenticity (2022-01-26). C2PA Releases Specification of World’s First Industry Standard for Content Provenance. C2PA. https://c2pa.org/c2pa-releases-specification-of-worlds-first-industry-standard-for-content-provenance/.
  10. ^a ^b Ryan-Mosley, Tate (2023-07-28). Cryptography may offer a solution to the massive AI-labeling problem. MIT Technology Review. https://www.technologyreview.com/2023/07/28/1076843/cryptography-ai-labeling-problem-c2pa-provenance/.
  11. ^ Kirchenbauer, John; Geiping, Jonas; Wen, Yuxin; Katz, Jonathan; et al. (2023). A Watermark for Large Language Models. Proceedings of the 40th International Conference on Machine Learning (ICML 2023). https://doi.org/10.48550/arXiv.2301.10226 https://arxiv.org/abs/2301.10226.
  12. ^ Dathathri, Sumanth; See, Abigail; Ghaisas, Sumedh; Huang, Po-Sen; et al. (2024-10). Scalable watermarking for identifying large language model outputs. Nature. https://doi.org/10.1038/s41586-024-08025-4 https://www.nature.com/articles/s41586-024-08025-4.
  13. ^a ^b European Commission (2026). Code of Practice on Transparency of AI-generated Content. Shaping Europe’s digital future. https://digital-strategy.ec.europa.eu/en/policies/code-practice-ai-generated-content.
  14. ^ Paul, Weiss, Rifkind, Wharton & Garrison LLP (2026-08-04). EU Finalises Transparency Rules for AI-Generated Content. https://www.paulweiss.com/insights/client-memos/eu-finalises-transparency-rules-for-ai-generated-content.
Available in