Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

content provenance and the generated web (authored by agents unless marked 🧑)

  • reviews tracing content origin and measuring generated content

    • each topic includes possible research experiments
  • generated-web measurement

    • prevalence studies and their relationship to DeGenTWeb
  • effects_of_generated_content: what AI-generated content has been measured to do, separated from what is only argued.

  • text_watermarking: watermarks in LLM text, how they break, and what Google, Anthropic and OpenAI deploy.

  • image_watermarking: watermarks in AI-generated images, audio and video.

  • c2pa: C2PA and other cryptographic provenance, its security analyses, and who has adopted it.

  • camera_authentication: cameras and phones that sign photos, proofs of edits, and photographing a screen.

  • labeling_rules_and_practice: laws and platform practice for labeling AI-generated content, and audits of them.

research priorities, recommended by agents

  • first: follow signed images through upload and republication
    • controlled experiment
    • separates lost credentials, invalid signatures, and untrusted signers
    • builds on the human’s existing photo authentication work
  • second: measure visible provenance in a defined web sample
    • report C2PA credentials, recoverable manifests, and detector-accessible watermarks separately
    • absence of a mark does not establish human authorship
    • pilot image collection and validator agreement before scaling
  • third: compare page, site, and text-volume estimates on the same sample
  • later: estimate which generated pages people actually read
    • requires credible audience data and a clear sampling frame
    • traffic estimates alone cannot establish that generated pages caused other sites to lose readers

remaining scope

  • these rankings are agent recommendations, not established novelty claims
  • detailed files preserve useful drafts and state their own unread sources and limits
  • verify deployment and legal claims against current primary sources before an experiment
  • the existing Extra High consultation addresses text detection
    • it does not independently review every provenance proposal
  • the camera review compares four edit-proof systems
    • independent evaluation of deployed depth-based copy detection remains future work
  • recovered Extra High advice and assessment
    • recommends signed-photo publication paths
    • the answer acted on the supplied requirements as a task
      • its ranked opinions and controls were assessed independently

Last edited: