Skip to content

AI Systems Handbook / Chapter 45

Intellectual Property, Content, and Provenance

Govern generative content from input rights and confidential prompts through human review, provenance, publication, and downstream correction.

The Image Nobody Could Clear

A marketing team publishes a generated image of a runner wearing its new shoe. A photographer then claims that the pose and setting reproduce one of her campaign photographs. The team opens the final file and finds no useful history. Nobody can identify the reference assets, prompt, model version, editor, license terms, or whether the runner depicts a real person. Ordinary metadata disappeared during export.

The complaint may prove well founded or mistaken. Either way, the company cannot clear the image, defend its process, or reliably find every version in circulation. The failure began before generation and continued after publication.

Content governance is a chain of custody. Control the rights and confidentiality of inputs, the permitted use and review of outputs, the record of human and machine contributions, the provenance presented downstream, and the ability to correct or withdraw.

A content chain of custody moves from authorized inputs through recorded generation, rights and harm review, provenance binding, approved publication, and downstream correction, while an evidence ledger runs beneath every stage.
Provenance begins before a prompt and continues after publication. A durable evidence ledger connects source rights, generation, review, disclosure, and correction.

Start the Record Before the Prompt

The team first needs to recover what entered the system. “We own the reference image” is too coarse: an organization may own a photograph while a license, model release, location agreement, trademark, or contract limits adaptation and commercial use. A licensed asset may permit publication but prohibit model input. Customer material may arrive with promises about purpose, isolation, retention, or deletion. Public access to an image establishes neither permission nor reliability.

Before a reference reaches a model, give it an asset identifier and record its source, custodian, authority or license, permitted purpose, territory, medium, confidentiality class, consent where relevant, and expiry. Record whose likeness, voice, marks, or other protected interests it contains. The useful question is not simply Can we access this? It is What exactly are we authorized to do with it here?

That question follows the material into generation. Prompts can contain trade secrets, unpublished work, source code, personal data, legal strategy, or customer material. Prompt text, retrieved context, uploads, outputs, feedback, and review queues need the same decisions about access, purpose, retention, deletion, and incident response as other content systems. The provider’s terms also matter: what it may retain or use, which model version produced the output, and which restrictions attach downstream.

Do not make a designer solve this alone at the moment of prompting. Approved environments, registered source connectors, blocked data classes, role-based templates, and data-loss controls turn policy into a usable path. When authority is absent or ambiguous, the input stops there or enters a named exception review.

Let Exposure Determine the Review

The runner image now has a source record and a generation record. It still is not cleared for a global campaign. “Who owns the output?” compresses several questions that may have different answers across law, contract, territory, and use:

Does applicable law recognize protectable human authorship in the work? Does the image reproduce or closely resemble protected material? Do source licenses, provider terms, or customer contracts restrict this use? Does it contain a person’s likeness, a trademark, trade dress, or confidential information? Could it deceive, defame, discriminate, or create a dangerous product claim? Does the channel require attribution, consent, or disclosure?

Similarity search, source lookup, moderation, and provenance tools provide evidence; none supplies a universal clearance decision. Qualified review is needed when the answer depends on jurisdiction or disputed rights. The U.S. Copyright Office, for example, treats copyrightability of generative output as a question of human authorship and control over expressive elements. That analysis is useful in its jurisdiction, not a worldwide ownership rule.

Preserve the work people actually did—selection, arrangement, revision, factual verification, creative decisions, and approval. Do not reconstruct a flattering human-authorship story after a dispute.

The depth of review follows exposure. Private ideation with approved, non-sensitive inputs may need only a recorded tool and source boundary. Routine internal material can use standard review. Commercial publication, brand claims, licensed sources, realistic people, political or public-interest content, children, health, finance, safety, or wide distribution need enhanced factual, rights, safety, brand, privacy, and accessibility review. Deceptive impersonation, non-consensual intimate content, unauthorized confidential data, evasion of rights controls, and uses barred by law, contract, or organizational policy do not enter an ordinary approval queue.

The shoe campaign therefore takes the enhanced route. The reviewers can see the source assets and terms; they have criteria, time, and authority to reject. “Human reviewed” without those conditions is only a description of traffic through an inbox.

Bind the Claim to the Asset

Once approved, the image receives a durable asset identifier. Its evidence ledger links the source rights and custodians; model, tool, version, account, and generation time; prompt or a protected reference to it; edits and contributors; factual, rights, safety, brand, and accessibility reviews; approval and publication channels; and any later corrections, withdrawals, derivatives, or downstream notices.

The ledger is the organization’s private evidence. A credential attached to the published asset makes selected claims portable. The Coalition for Content Provenance and Authenticity (C2PA) specification defines Content Credentials as cryptographically bound provenance structures. A recipient can inspect who signed the recorded assertions and whether the credential or bound asset was subsequently altered.

This answers a narrow and valuable question about provenance. It does not prove that every assertion is complete, that the depicted event occurred, that the source was licensed, or that the content is lawful or harmless. A valid signature can bind a false claim; an unsigned asset can be authentic. Trust still depends on the evidence, the signer, and the use.

The team also assumes that binding will break. A platform may strip metadata; a screenshot or crop may sever the hard binding; an unsigned copy may circulate. The server-side ledger remains authoritative, and high-value assets can use durable credentials, content fingerprints, or a public verification route to reconnect a copy with its record. Public assertions should not expose confidential prompts, personal data, sensitive locations, or vulnerable creators merely to make provenance exhaustive.

Make Disclosure Useful at the Moment of Use

The campaign team must now decide what the audience needs to know. A visible label may be important when a realistic person or event could be mistaken for documentary evidence. A translated or lightly assisted product description may call for different wording from a fabricated voice or materially manipulated news image. Machine-readable marking helps systems carry context; visible language helps people interpret it.

Place disclosure where the user encounters the risk: publication, download, remix, or redistribution. Say whether content was generated, materially manipulated, translated, summarized, or assisted when that distinction helps. Provide source, method, limitation, or verification detail when a person can use it. Make the treatment accessible and localized, then test whether people understand it. An unfamiliar provenance icon does not become meaningful through repetition alone.

Rules are changing and differ by role and jurisdiction. In the European Union, Article 50 transparency obligations are scheduled to apply from 2 August 2026, and a voluntary code published in June 2026 offers practices for marking and labeling certain generated or manipulated content. Teams need qualified review of scope; neither a universal “AI-made” badge nor silence is a durable compliance strategy.

A Complaint Tests the Whole Chain

The photographer’s complaint should reach the asset record, not a generic support queue. The response owner can pause paid distribution, locate known derivatives and channels, preserve the disputed files and decisions, notify partners, and place a corrected or replacement asset. The rights reviewer can compare the registered references and generation history with the claimed photograph. If the complaint reveals a missed source or unsafe similarity, the team can block that input, revise review criteria, and add a regression case.

That response depends on every earlier control: permitted sources, protected accounts and prompts, constrained tools, output review, durable provenance, a distribution inventory, and named withdrawal authority. Moderation at generation time cannot replace the chain. Neither can a final visual inspection.

Operational measures should reveal whether the chain works: severe misses, reviewer disagreement, escalation, correction time, credential coverage and validation, repeated claims, and whether labels survive downstream. A high block rate may reflect attack pressure, overblocking, or a badly designed route. Investigate the cause rather than reward the number.

Decide the Rule for the Next Campaign

The team cannot eliminate every disputed right or misleading interpretation. It can decide which content may be created, what evidence must travel with it, who may approve its use, and how it will respond when the decision is challenged. Put those decisions in an AI Content Use Policy that people can operate:

  • Scope and purpose: covered people, tools, models, content, channels, regions, third parties, approved uses, audiences, exposure routes, and non-AI alternatives.
  • Inputs and generation: authority or license, confidentiality, privacy, consent, provenance, prohibited sources, approved environment, model version, account, prompt handling, retrieval, safety settings, and retention.
  • Review and publication: factual, source, rights, privacy, safety, fairness, brand, accessibility, and domain criteria; reviewer authority; permitted output use; attribution; human contribution; disclosure; and downstream restrictions.
  • Provenance and response: asset identifier, private evidence ledger, credentials, validation, privacy limits, distribution record, complaints, claims, correction, withdrawal, notice, and evidence preservation.
  • Assurance: training, sampling, operating measures, audits, vendor changes, exceptions, policy updates, and expiry.

Apply the policy to three proposed outputs: an internal mood board made from approved stock assets, the global runner campaign, and a realistic video of the chief executive announcing a product recall. For each, name the source evidence, review route, disclosure, provenance, withdrawal owner, and any reason not to proceed. If the same answer fits all three, the policy has not yet captured exposure.

The goal is not risk-free content. It is a content operation that can explain what it made, why it was allowed, who decided, and what happens next when the decision is disputed.

Source Notes