FROM AGPEDIA — AGENCY THROUGH KNOWLEDGE

Model provenance

Model provenance is the documented origin and history of an artificial intelligence model: what data it was trained on, which earlier models it was built from, how it was trained and modified, and who did that work. The term also covers ways to check that history, such as cryptographic signatures that show a model file has not been altered, and statistical tests that estimate whether one model was derived from another. Security engineers treat a model's provenance as part of the AI supply chain, because a model's behavior depends on inputs that people using the model usually cannot see.[1][2]

Interest in model provenance grew as foundation models were increasingly adapted and redistributed by parties other than their original developers. Audits have found that the provenance of training data is often poorly documented: a 2024 study of more than 1,800 text datasets found license information missing for more than 70% of datasets on popular hosting sites, and wrong for more than 50%.[3:1] Since 2025, laws in the European Union and California have required some AI developers to publish summaries of their training data.[4][5]

Terminology and scope

Model provenance is distinct from AI content provenance, which concerns whether a particular image, video, or text was made by AI. Model provenance concerns the history of the AI system itself. It overlaps with data provenance, a broader field that tracks the origin and processing of data in databases and scientific workflows; for AI models, the provenance of training data is one component of model provenance.

The term is used in two related senses. In a documentation sense, provenance is a record that developers publish or attach to a model, such as a model card or an AI bill of materials. In a forensic sense, provenance is something investigators infer from the model itself when records are missing or untrusted. Ivica Nikolić, Teodora Baluta, and Prateek Saxena frame the forensic question as "whether one model is derived from another," and argue that "tracking model origins is crucial both for protecting intellectual property and for identifying derived models when biases or vulnerabilities are discovered in foundation models."[2]

Components

Component What provenance records Example methods
Training data Sources, licenses, collection dates, processing, presence of personal or copyrighted material Datasheets, dataset audits, training-data summaries required by law
Model lineage Base models, fine-tuning steps, merges, and other derivations Model cards, AI bills of materials, provenance testing
Model artifacts That a downloaded model file is the one its developer released Cryptographic hashes and signatures

Training data

Many AI models are trained on collections that combine thousands of smaller datasets, each with its own source and license. Shayne Longpre and colleagues at the Data Provenance Initiative traced the lineage of more than 1,800 text datasets used to fine-tune language models, including their sources, creators, and licenses. They found frequent errors on dataset hosting sites, with license omission above 70% and error rates above 50%, and described the result as "a crisis in misattribution and informed use of popular datasets."[3:1]

Provenance also includes whether the owners of source material consented to its use. In a 2024 audit of about 14,000 web domains, Longpre and colleagues found that in a single year, restrictions added by websites had made about 5% of all tokens in the widely used C4 web dataset, and more than 28% of its most actively maintained sources, fully restricted from AI use. They also found inconsistencies between what websites said in their terms of service and what they signaled in their robots.txt files, the standard way websites tell automated crawlers what they may access.[6]

Model lineage

Developers frequently build new models by fine-tuning an existing one on additional data. A fine-tuned model inherits properties of its base model, including its licensing terms and any flaws discovered later. Nikolić and colleagues note that the growth of fine-tuning creates "challenges in enforcing licensing terms and managing downstream impacts."[2]

Model artifacts

A trained model is distributed as files containing its parameters, often through public model hubs. Those files can be tampered with. In February 2024, researchers at the security company JFrog reported finding about 100 models on the Hugging Face hub with malicious functionality. One technique abused Python's pickle file format to run arbitrary code when a model file was loaded; one such model contained a payload that could open a remote "reverse shell," giving an attacker control of the victim's computer.[7]

Documentation practices

In 2019, Margaret Mitchell and colleagues proposed model cards: "short documents accompanying trained machine learning models" that report how a model performs across different conditions and groups, along with its intended uses and how it was evaluated.[8] A companion proposal, datasheets for datasets by Timnit Gebru and colleagues, borrowed the practice from the electronics industry, in which every component comes with a datasheet describing its characteristics. They proposed that each dataset document its "motivation, composition, collection process, recommended uses, and so on."[9]

AI bills of materials extend the software industry's software bill of materials (SBOM) to AI. An SBOM is a machine-readable list of the components that make up a piece of software. Version 3.0 of the SPDX standard, released on April 16, 2024, added profiles for describing AI models and datasets, including "AI model training and characterization" and "data set provenance." SPDX is published as the international standard ISO/IEC 5962:2021.[10]

Stanford's 2023 Foundation Model Transparency Index scored 10 major developers on 100 transparency indicators. It found that disclosure was worst for "upstream" matters such as data, labor, and compute, and that no company scored points on indicators for data creators, the copyright and license status of data, or copyright mitigations. The highest overall score was 54 out of 100 (Meta) and the mean was 37.[11:1]

Integrity and security

Provenance records matter for security because both training data and model files can be manipulated. Nicholas Carlini and colleagues showed in 2023 that some web-scale datasets, which are distributed as lists of links rather than copies of the content, can be poisoned by buying expired domains that the links point to and replacing the content. They estimated that an attacker could have poisoned 0.01% of the LAION-400M or COYO-700M datasets for about $60, and proposed defenses including distributing cryptographic hashes of all content, so that people downloading a dataset can check they receive the same data its creators indexed.[12]

Model signing applies the same idea to model files. In April 2025, the Open Source Security Foundation (OpenSSF) released version 1.0 of a model signing library and command-line tool, developed with contributors from Google, NVIDIA, and HiddenLayer. It lets developers sign models of any format and size so that users can verify a model has not changed since release. The project's authors note that "the team that trains a foundation model is not the same as the team that deploys the model into production," which makes integrity checks between those teams necessary.[1]

Provenance testing

When documentation is missing or disputed, researchers have developed methods to infer lineage from a model's behavior. Nikolić, Baluta, and Saxena proposed a statistical test that compares the outputs of two models against a baseline of unrelated models, using only black-box access (querying the models without seeing their internals). Tested on more than 600 models ranging from 30 million to 4 billion parameters, it identified derived models with 90–95% precision and 80–90% recall.[2]

Regulation

Jurisdiction Measure Main requirement In force
European Union AI Act, Article 53 Providers of general-purpose AI models must keep technical documentation, adopt a copyright policy, and publish a summary of training content using a Commission template August 2, 2025 (August 2, 2027, for models already on the market)
California AB 2013, Generative Artificial Intelligence Training Data Transparency Act Developers of public generative AI systems must post a high-level summary of their training datasets covering 12 listed items January 1, 2026

European Union

Article 53 of the EU AI Act requires providers of general-purpose AI models to keep technical documentation "including its training and testing process and the results of its evaluation," to adopt a policy for complying with EU copyright law, and to "make publicly available a sufficiently detailed summary about the content used for training." Providers of free and open-source models are exempt from some of these duties unless their models are designated as posing systemic risk.[4] The European Commission published the required template for the training-content summary on July 24, 2025. It asks for information such as the main datasets used, a list of domain names scraped, and the categories of data sources.[13] The obligations apply from August 2, 2025; models placed on the market before that date have until August 2, 2027, to comply.[4][14]

California

California's AB 2013, enacted in 2024, requires developers of generative AI systems made available to the public in California to post documentation about their training data by January 1, 2026, and before each later release. It covers systems released on or after January 1, 2022. The documentation must include a "high-level summary" covering 12 items, including the sources or owners of the datasets, whether they include copyrighted material or personal information, whether they were purchased or licensed, the period in which the data were collected, and whether synthetic data were used.[5]

Limitations

Documentation depends on developers' willingness and ability to disclose. The Foundation Model Transparency Index authors link the absence of disclosure about data creators and copyright to "ongoing litigation" over copyright and intellectual property, and note that methods for identifying the creators of web-scale scraped data are still immature.[11:1]

Forensic methods have limits too. The provenance test by Nikolić and colleagues identified 80–90% of derived models in its benchmark, which means some derived models went undetected.[2] Signatures show that a file has not changed since it was signed, but not that the signer's claims about training data are true.[1]

Analysis

Model provenance affects the agency of several groups. Developers who build on existing models need provenance to know what licensing terms, flaws, and risks they inherit. People whose writing, images, or personal data are used for training need it to know whether and how their work was used, and to exercise any rights they have. Users and regulators need it to judge whether a model is suitable for a given purpose. The audits cited above suggest that, as of 2023–2024, much of this information was missing or unreliable, which limits all three groups' ability to make informed choices.[3:1][11:1]

The training-data summaries now required in the EU and California create a baseline of disclosure, but they are summaries rather than full records. Whether they are detailed enough to let rights holders and researchers check specific claims could be tested by comparing published summaries against independent audits of the same models.

  1. ^a ^b ^c Maruseac, Mihai; Sablotny, Martin; Wickens, Eoin; Major, Daniel (2025-04-04). Launch of Model Signing v1.0: OpenSSF AI/ML Working Group Secures the Machine Learning Supply Chain. Open Source Security Foundation Blog. https://openssf.org/blog/2025/04/04/launch-of-model-signing-v1-0-openssf-ai-ml-working-group-secures-the-machine-learning-supply-chain/.
  2. ^a ^b ^c ^d ^e Nikolić, Ivica; Baluta, Teodora; Saxena, Prateek (2025-02-02). Model Provenance Testing for Large Language Models. arXiv. https://doi.org/10.48550/arXiv.2502.00706 https://arxiv.org/abs/2502.00706.
  3. ^a ^b ^c ↗ license-errors Longpre, Shayne; Mahari, Robert; Chen, Anthony; Obeng-Marnu, Naana; et al. (2024-08-30). A large-scale audit of dataset licensing and attribution in AI. Nature Machine Intelligence. https://doi.org/10.1038/s42256-024-00878-8 https://www.nature.com/articles/s42256-024-00878-8.
  4. ^a ^b ^c European Parliament and Council of the European Union (2024-06-13). Regulation (EU) 2024/1689 (Artificial Intelligence Act), Article 53: Obligations for Providers of General-Purpose AI Models. https://artificialintelligenceact.eu/article/53/.
  5. ^a ^b California State Legislature (2024-09-28). AB-2013 Generative artificial intelligence: training data transparency. https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013.
  6. ^ Longpre, Shayne; Mahari, Robert; Lee, Ariel; Lund, Campbell; et al. (2024-07). Consent in Crisis: The Rapid Decline of the AI Data Commons. arXiv. https://doi.org/10.48550/arXiv.2407.14933 https://arxiv.org/abs/2407.14933.
  7. ^ Toulas, Bill (2024-02-28). Malicious AI models on Hugging Face backdoor users’ machines. BleepingComputer. https://www.bleepingcomputer.com/news/security/malicious-ai-models-on-hugging-face-backdoor-users-machines/.
  8. ^ Mitchell, Margaret; Wu, Simone; Zaldivar, Andrew; Barnes, Parker; et al. (2019-01). Model Cards for Model Reporting. Proceedings of the Conference on Fairness, Accountability, and Transparency (FAT* ’19). https://doi.org/10.1145/3287560.3287596 https://arxiv.org/abs/1810.03993.
  9. ^ Gebru, Timnit; Morgenstern, Jamie; Vecchione, Briana; Vaughan, Jennifer Wortman; et al. (2021-12). Datasheets for Datasets. Communications of the ACM. https://arxiv.org/abs/1803.09010.
  10. ^ Linux Foundation (2024-04-16). SPDX 3.0 Revolutionizes Software Management in Systems with Enhanced Functionality and Streamlined Use Cases. https://www.linuxfoundation.org/press/spdx-3-revolutionizes-software-management-in-systems-with-enhanced-functionality-and-streamlined-use-cases.
  11. ^a ^b ^c ↗ upstream-opacity Bommasani, Rishi; Klyman, Kevin; Longpre, Shayne; Kapoor, Sayash; et al. (2023-10-19). The Foundation Model Transparency Index. arXiv. Stanford Center for Research on Foundation Models. https://doi.org/10.48550/arXiv.2310.12941 https://arxiv.org/abs/2310.12941.
  12. ^ Carlini, Nicholas; Jagielski, Matthew; Choquette-Choo, Christopher A.; Paleka, Daniel; et al. (2023-02-20). Poisoning Web-Scale Training Datasets is Practical. arXiv. https://doi.org/10.48550/arXiv.2302.10149 https://arxiv.org/abs/2302.10149.
  13. ^ European Commission (2025-07-24). Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models. Shaping Europe’s digital future. https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models.
  14. ^ European Parliament and Council of the European Union (2024-06-13). Regulation (EU) 2024/1689 (Artificial Intelligence Act), Article 111: AI Systems Already Placed on the Market or Put into Service and General-Purpose AI Models Already Placed on the Marked. https://artificialintelligenceact.eu/article/111/.
Available in