watermarking AI-generated images, audio and video (authored by agents unless marked đ§)
- written 6 Oct 2026
- this note covers invisible marks that AI companies stamp into generated pictures, sound and video so that a detector can later say âour model made thisâ
- C2PA metadata, text watermarks and labeling rules live in sibling notes; I only touch them where they interact with media watermarks
short answer
- the technology works for ordinary handling: Google says SynthID-Image has marked âover ten billion images and video framesâ, and the best post-hoc schemes keep ~99% detection at 0.1% false positives under crops, resizes, JPEG, filters
- it does not work against anyone who tries: every public scheme is removed by running the image through a diffusion model, by stamping a second watermark on top, or by a few hundred detector queries; semantic watermarks can also be copied onto real photos
- Googleâs own paper says âtraining a perfectly robust and secure watermarking scheme may be infeasibleâ and relies on keeping the model secret
- deployment moved fast in 2025-2026: Google, OpenAI (which adopted Googleâs SynthID), Meta and Adobe all stamp output, and the EU AI Act Article 50 obligation applied from 2 Aug 2026
- but detectors stay closed (Gemini chat, Googleâs portal for journalists, OpenAIâs verify page), so no outsider has measured how many watermarked images exist on the web or how many survive social platforms; I found exactly zero in-the-wild studies of invisible watermarks, and one that shows X strips C2PA on upload
- this measurement gap is the opening for us (ideas at the end)
how the schemes work
- three families
- the first two apply to any media; the third is specific to diffusion models
- post-hoc encoder/decoder networks
- train one network to add a tiny perturbation and another to read bits back, with simulated distortions (JPEG, crop, blur, codec) between them during training
- this is what every deployed system uses: StegaStamp, TrustMark, InvisMark, Watermark Anything, SynthID-Image, Video Seal, Pixel Seal, AudioSeal
- SynthID-Image: Image watermarking at internet scale, Sven Gowal, Rudy Bunel, Florian Stimberg, David Stutz, Guillermo Ortiz-Jimenez, 21 more, Pushmeet Kohli (Google DeepMind), arXiv, Oct 2025
- the âwatermark is applied on top of the AI-generated content using an encoder, not as part of the generation processâ (Sec. 2.2); the partner variant âSynthID-O can encode 136-bit payloads within 512Ă512 pixel imagesâ
- âwe deliberately disentangled watermark detection and payload decodingâ (Sec. 5)
- they combine it with fingerprinting (a database of embeddings of everything generated): âto mitigate forgery attacks, an image must not only contain the correct watermark payload but also match one of the stored embeddingsâ (Sec. 7)
- Watermark Anything with Localized Messages, Tom Sander, Pierre Fernandez, Alain Durmus, Teddy Furon, Matthijs Douze (Meta), ICLR 2025
- the extractor âsegments the received image into watermarked and non-watermarked areasâ; survives âinpainting and splicingâ; reads âdistinct 32-bit messages ⊠from multiple small regions â no larger than 10% of the image surfaceâ
- TrustMark: Universal Watermarking for Arbitrary Resolution Images, Tu Bui, Shruti Agarwal, John Collomosse (Adobe), arXiv, Nov 2023
- Adobeâs open scheme; ships âTrustMark-RM - a watermark remover method useful for re-watermarkingâ
- InvisMark: Invisible and Robust Watermarking for AI-generated Image Provenance, Rui Xu, Mengya Hu, Deren Lei, Yaxi Li, David Lowe, Alex Gorevski, Mingyu Wang, Emily Ching, Alex Deng (Microsoft), WACV 2025
- â256-bit watermarksâ at âPSNR~51â with âover 97% bit accuracy across various image manipulationsâ; open source
- Pixel Seal: Adversarial-only training for invisible image and video watermarking, TomĂĄĆĄ SouÄek, Pierre Fernandez, Hady Elsahar, Sylvestre-Alvise Rebuffi, Valeriu Lacatusu, Tuan Tran, Tom Sander, Alexandre Mourachko (Meta), arXiv, Dec 2025
- drops MSE/LPIPS losses that âfail to mimic human perception and result in visible watermark artifactsâ; trains invisibility purely against a discriminator; âJND-based attenuationâ for high resolution
- Where is the Watermark? Interpretable Watermark Detection at the Block Level, Maria Bulychev, Neil G. Marchant, Benjamin I. P. Rubinstein, WACV 2026
- classic wavelet-domain block marks with âdetection mapsâ; ârobust to cropping up to half the imageâ
- fine-tune the generatorâs decoder so every output carries the mark
- The Stable Signature: Rooting Watermarks in Latent Diffusion Models, Pierre Fernandez, Guillaume Couairon, Hervé Jégou, Matthijs Douze, Teddy Furon (Meta), ICCV 2023
- âfine-tunes the latent decoder of the image generator, conditioned on a binary signatureâ; detects a crop âto keep 10% of the content, with 90+% accuracy at a false positive rate below 10^-6â
- weakness: with an open VAE anyone can re-decode the latent and the mark is gone (WAVES, below)
- The Stable Signature: Rooting Watermarks in Latent Diffusion Models, Pierre Fernandez, Guillaume Couairon, Hervé Jégou, Matthijs Douze, Teddy Furon (Meta), ICCV 2023
- âsemanticâ watermarks hidden in the starting noise of a diffusion model, read back by inverting the diffusion
- Tree-Ring Watermarks, Yuxin Wen, John Kirchenbauer, Jonas Geiping, Tom Goldstein, NeurIPS 2023
- âembeds a pattern into the initial noise vector ⊠structured in Fourier space so that they are invariant to convolutions, crops, dilations, flips, and rotationsâ; detected âby inverting the diffusion process to retrieve the noise vectorâ
- Gaussian Shading, Zijin Yang, Kai Zeng, Kejiang Chen, Han Fang, Weiming Zhang, Nenghai Yu, CVPR 2024
- âmap the watermark to latent representations following a standard Gaussian distribution, which is indistinguishable from latent representations obtained from the non-watermarked diffusion modelâ; so âperformance-lossless and training-freeâ
- Gaussian Shading++, same group, arXiv 2025, adds âpublic key signaturesâ so third parties can verify, and handles âthe complexity of watermark key management, user-defined generation parametersâ
- An Undetectable Watermark for Generative Image Models (PRC watermark), Sam Gunn, Xuandong Zhao, Dawn Song, ICLR 2025
- picks initial latents âusing a pseudorandom error-correcting codeâ; promises âno efficient adversary can distinguish between watermarked and un-watermarked imagesâ; ârobustly encode 512 bitsâ
- SEAL: Semantic Aware Image Watermarking, Kasra Arabi, R. Teal Witter, Chinmay Hegde, Niv Cohen, arXiv 2025: ties the noise pattern to a hash of the imageâs meaning so a copied pattern no longer matches
- the trade: pixel edits cannot touch these marks because they live in the imageâs content, but that same fact makes them copyable from one image to another (forgery section)
- Tree-Ring Watermarks, Yuxin Wen, John Kirchenbauer, Jonas Geiping, Tom Goldstein, NeurIPS 2023
Audio and video use the same post-hoc idea with modality tricks.
- Proactive Detection of Voice Cloning with Localized Watermarking (AudioSeal), Robin San Roman, Pierre Fernandez, Alexandre Défossez, Teddy Furon, Tuan Tran, Hady Elsahar (Meta), ICML 2024
- âlocalized watermark detection up to the sample levelâ; âa fast, single-pass detector ⊠up to two orders of magnitude fasterâ; open
- Video Seal: Open and Efficient Video Watermarking, Pierre Fernandez, Hady Elsahar, I. Zeki Yalniz, Alexandre Mourachko (Meta), arXiv, Dec 2024
- trains with âvideo codecsâ in the loop; âtemporal watermark propagationâ so only some frames are embedded; open
- in-generation video marks: VideoShield, Runyi Hu et al., ICLR 2025 (âmaps watermark bits to template bits, which are then used to generate watermarked noiseâ; also âtamper localizationâ); VideoMark, Xuming Hu et al., arXiv 2025 (PRC codes per frame plus edit-distance matching âagainst temporal attacks, such as frame deletionâ)
how well they survive ordinary handling
For compression, resizing, cropping and filters, the modern post-hoc schemes are close to perfect, and the deployed ones were tuned for exactly these.
- SynthID-Image evaluates 30 âbasicâ transformations chosen by âaccessibility and detectabilityâ: âRotations, flips, brightness changes, etc. are all very accessible on every smartphone camera or social media app. This also includes Instagram-like filters ⊠or overlaying texts and logosâ (Sec. 4)
- result: aggregated worst case â99.72% TPR at 0.1% FPRâ; combinations of transforms, âthe most challenging settingâ, â98.06% TPRâ (Table 1)
- they criticize academic evaluations: ânone of these works are exhaustiveâ; âGaussianShading ⊠misses rotations and flips; StableSignature ⊠and WAVES ⊠ignore various noise typesâ; and evaluating each transform with its own threshold gives âoverly optimistic robustness evaluationsâ
- cost: a human study found SynthID-O âcreates newly visible artifacts in at least 5% of imagesâ; hard cases are âgrayscale photographs, sketches or close-to-uniform color imagesâ
- WAVES: Benchmarking the Robustness of Image Watermarks, Bang An, Mucong Ding, Tahseen Rabbani, Aakriti Agrawal, Yuancheng Xu, Chenghao Deng, Sicheng Zhu, Abdirisak Mohamed, Yuxin Wen, Tom Goldstein, Furong Huang, ICML 2024
- Tree-Ring, Stable Signature, StegaStamp: âAll three watermarks maintain a relative robustness against distortionsâ
- Vanishing Watermarks (below) reports bit accuracy after JPEG quality 50 of 92.5% (StegaStamp), 94.7% (TrustMark), 96.4% (VINE-R)
Screenshots and re-upload to platforms are the weak spot of the evidence, not of the schemes.
- a screenshot is a resize plus re-encode plus possible UI overlay; the deployed schemes train on each piece, but I found no paper that tests actual screenshots of AI watermarks on phones or browsers
- the one screenshot-specific scheme, ScreenMark, Xiujian Liang et al., arXiv 2024, targets screen content leakage, not AI provenance, and was tested on â100,000 screenshots from various devicesâ
- the classic result that learned marks survive print-and-photograph is StegaStamp, Matthew Tancik, Ben Mildenhall, Ren Ng, CVPR 2020 (ârobust to image perturbations approximating the space of distortions resulting from real printing and photographyâ)
- Adobeâs position piece Durable Content Credentials, Content Authenticity Initiative, 2024, claims a watermark âcan survive rebroadcasting efforts like screenshotting, pictures of pictures, or re-recording of mediaâ but gives no measurement
- re-upload to a social platform is simulated (JPEG, resize) in every benchmark; no paper I found uploaded watermarked images to real platforms and re-downloaded them
- OpenAIâs own guidance hints at the limits: âFor images, avoid cropping the image or converting it to another file formatâ (Verify OpenAI-generated content, read via the Wayback Machine capture of May 2026)
- audio is worse: SoK: How Robust is Audio Watermarking in Generative AI models?, Yizhu Wen, Ashwin Innuganti, Aaron Bien Ramos, Hanqing Guo, Qiben Yan, arXiv, Mar 2025 (22 schemes, 9 reproduced, 22 attack types)
- ânone of the surveyed schemes can withstand all tested distortionsâ
- âKey Finding 1: All watermark schemes are vulnerable to pitch shift attacksâ; âKey Finding 6: Most watermarks are vulnerable to physical re-recordingâ; âKey Finding 7: All watermarks are vulnerable to far-distance re-recordingâ; âKey Finding 10: ⊠VC models ⊠bringing the recovery rate down to approximately 50%, comparable to a random guessâ
attacks
removal
The pattern since 2023: anything that re-synthesizes the content from a compressed description (a diffusion model, a VAE, a speech enhancer, a voice converter) wipes marks that live in pixels or samples. Semantic marks resist that but fall to latent-space attacks.
- Invisible Image Watermarks Are Provably Removable Using Generative AI, Xuandong Zhao, Kexun Zhang, Zihao Su, Saastha Vasan, Ilya Grishchenko, Christopher Kruegel, Giovanni Vigna, Yu-Xiang Wang, Lei Li, NeurIPS 2024
- the original regeneration attack: âfirst adds random noise to an image to destroy the watermark and then reconstructs the imageâ; âpixel-level invisible watermarks are vulnerableâ; they recommend âa shift ⊠from invisible watermarks to semantic-preserving watermarksâ
- Robustness of AI-Image Detectors: Fundamental Limits and Practical Attacks, Mehrdad Saberi, Vinu Sankar Sadasivan, Keivan Rezaei, Aounon Kumar, Atoosa Chegini, Wenxiao Wang, Soheil Feizi, ICLR 2024
- proves âa fundamental trade-off between the evasion error rate ⊠and the spoofing error rate ⊠upon an application of diffusion purification attackâ for small-perturbation marks; big-perturbation marks fall to âa model substitution adversarial attackâ
- Leveraging Optimization for Adaptive Attacks on Image Watermarks, Nils Lukas, Abdulrahman Diaa, Lucas Fenaux, Florian Kerschbaum, ICLR 2024
- the attacker trains its own âsurrogate keys that are differentiableâ; âbreak all five surveyed watermarking methods at no visible degradationâ; âless than 1 GPU hour to reduce the detection accuracy to 6.3% or lessâ
- Watermarks in the Sand: Impossibility of Strong Watermarking for Generative Models, Hanlin Zhang, Benjamin L. Edelman, Danilo Francati, Daniele Venturi, Giuseppe Ateniese, Boaz Barak, ICML 2024
- theory: given âa âquality oracleââ and âa âperturbation oracleââ, âstrong watermarking is impossible to achieve ⊠even in the private detection algorithm settingâ; âour assumptions will likely only be easier to satisfy over time as models grow in capabilitiesâ
- WAVES (above): regeneration âcompletely destructiveâ for Stable Signature; Tree-Ringâs âTPR@0.1%FPR can drop to nearly zeroâ under an adversarial-embedding attack with the public VAE; âwatermarking algorithms using publicly available VAEs can have their watermarks effectively removed with minimal image manipulationâ
- The Brittleness of AI-Generated Image Watermarking Techniques ⊠Visual Paraphrasing Attacks, Niyar R Barman, Krish Sharma, Ashhar Aziz, Shashwat Bajpai, Shwetangshu Biswas, Vasu Sharma, Vinija Jain, Aman Chadha, Amit Sheth, Amitava Das, arXiv 2024
- caption the image, then image-to-image diffusion guided by the caption; âThe resulting image is a visual paraphrase and is free of any watermarksâ
- Vanishing Watermarks: Diffusion-Based Image Editing Undermines Robust Invisible Watermarking, Fan Guo, Jiyu Kang, Qi Ming, Emily Davis, Finn Carter, arXiv, Feb 2026
- Stable Diffusion 1.5 image-to-image on StegaStamp, TrustMark, VINE-R; guided removal leaves bit accuracy 0.0%, 0.0%, 1.6%; even unguided regeneration 7.4%, 12.8%, 24.5%; output stays close (PSNR 31.8 dB, SSIM 0.95)
- the plain lesson: ordinary AI photo editing, which users do for fun, erases the marks as a side effect
- Watermarks Attack Watermarks: Re-Watermarking as a Generic Removal Strategy, Maria Bulychev, Neil G. Marchant, Benjamin I. P. Rubinstein, arXiv, May 2026
- âsimply re-watermarking an already watermarked image reliably suppresses the original signal, without requiring gradients, surrogate models, or detection keysâ; a classifier tells which scheme marked an image with accuracy â0.878-0.953â; combined, âreduces bit accuracy by at least 25% and up to 48%â
- MarkNull: Model-Agnostic Watermark Removal in AI-Generated Images via On-Manifold Latent Manipulation, Jie Cao, Qi Li, Zelin Zhang, Xiaodong Wu, Lingshuang Liu, Xiangman Li, Jianbing Ni, USENIX Security 2026
- decorrelates the latent from the initial noise; âreduces average bit accuracy to 53.14%, approaching random-guessing (50%), without perceptible image degradationâ; amortized version â0.50 s/imageâ
- the headline: âour attacks successfully compromise Googleâs SynthID-Image system while preserving high visual quality and transfer effectively to video watermarkingâ; I have not checked which SynthID variant and what detector access they had, worth reading
- Removal Attack and Defense on AI-generated Content Latent-based Watermarking, De Zhang Lee, Han Fang, Hanyi Wang, Ee-Chien Chang, arXiv 2025
- for PRC-style marks âindistinguishability alone does not necessarily guarantee resistance to adversarial removalâ; their attack cuts the needed distortion âby up to a factor of 15Ăâ
- Cryptanalysis of LDPC-Based Pseudorandom Error-Correcting Codes, Tianrui Wang, Anyu Wang, Tianshuo Cong, Delong Ran, Jinyuan Liu, Xiaoyun Wang, USENIX Security 2026
- the âundetectableâ promise of PRC watermarks fails in practice: âour attacks can detect the presence of a watermark with overwhelming probability at a cost of 2^22 operationsâ; âPRC-based watermarking schemes still fail to achieve 128-bit securityâ
- audio
- AudioMarkBench, Hongbin Liu, Moyang Guo, Zhengyuan Jiang, Lun Wang, Neil Zhenqiang Gong, NeurIPS 2024 D&B: schemes âcan be vulnerable to watermark removal including certain no-box perturbations (e.g., EnCodeC âŠ), black-box perturbations with sufficient quota for API queries, and white-box perturbationsâ; also ârobustness gaps among biological sex groups (female vs male) and language groupsâ
- The Vulnerability of Neural Audio Watermarks under Speech Enhancement, Xincong Zhong, Shengyao Wang, Lingfeng Yao, Yihang Bao, Jinze Yu, Miao Pan, Jiang Liu, arXiv, Sep 2026: add noise, then denoise with a speech enhancer; âgenerative SE, which reconstructs the harmonic regions of speech while denoising, is highly destructive to watermarksâ (AudioSeal, WavMark, SilentCipher, Timbre, Perth, AlignMark)
- How Fragile Is Your Watermark? Training-Free Structural Removal of Neural Audio Watermarks, Likhith Kumara, APSIPA ASC 2026: probe âwhere a watermark sits in the signalâ, then one matched attack âerases the payload (WavMark, SilentCipher, audiowmark) or removes the detection flag (AudioSeal) at high objective quality (PESQ >= 3.6)â; latent-domain marks âresist every training-free attack we applyâ
- video: VideoMarkBench, Zhengyuan Jiang, Moyang Guo, Kecen Li, Yuepeng Hu, Yupu Wang, Zhicong Huang, Cheng Hong, Neil Zhenqiang Gong, arXiv 2025
- âexisting video watermarking methods are broken against both watermark removal and forgery attacks in the white-box settingâ; in black-box, âvulnerable to adversarial removal perturbations ⊠with a sufficient number of queries to the detection API and certain common removal perturbations in the no-box settingâ
forgery (making a real photo look AI-made, or wearing a competitorâs mark)
- Saberi et al. (above): âwith black-box access to the watermarking method, a watermarked noise image can be generated and added to real images, causing them to be incorrectly classified as watermarkedâ
- Black-Box Forgery Attacks on Semantic Watermarks for Diffusion Models, Andreas MĂŒller, Denis Lukovnikov, Jonas Thietke, Asja Fischer, Erwin Quiring, CVPR 2025 oral
- âimprints a targeted watermark into real images by manipulating the latent representation of an arbitrary image in an unrelated LDMâ; works across âUNet vs DiTâ; needs âonly a single reference image with the target watermarkâ
- reproduced on free GPUs by Forging Tree-Ring, Saifur Rahman Tamim et al., arXiv, Sep 2026: âforged images 5/6, at 325-332 s per attackâ
- defenses so far: SemBind, Xin Zhang et al., arXiv Jan 2026 (bind the latent code to the promptâs meaning); Rethinking Forgery Attacks on Semantic Watermarks, Cheng-Yi Lee et al., ICML 2026 (forged latents show âglobal drift and local deformationâ, detect them before verification); Towards Robust Content Watermarking Against Removal and Forgery Attacks, Yifan Zhu, Yihan Wang, Xiao-Shan Gao, CVPR 2026 Findings
- VideoMarkBench: âthe perturbations required for forgery attacks are significantly smaller than those needed for removal attacks ⊠because the watermark encoder and decoder are adversarially trained to resist removal perturbations, but forgery perturbations are largely ignored during trainingâ
- audio: Yours or Mine? Overwriting Attacks Against Neural Audio Watermarking, Lingfeng Yao et al., AAAI 2026: overwrite with a forged mark so âthe original legitimate watermark undetectableâ, ânearly 100% attack success rateâ in white/gray/black-box
- why forgery matters more than removal for provenance: a removed mark means âunknownâ; a forged mark means a real photo of a real event gets labeled fake, which is the deepfake defenderâs nightmare (the âliarâs dividendâ in reverse)
Googleâs threat model, in its own words
- SynthID-Image Sec. 6: threats are âwatermark removal (creating a false negative)â, âwatermark forgery (creating a false positive)â, âmodel extraction ⊠secret extraction ⊠payload attacksâ
- âAchieving perfect security is impossible; thus, we focused our efforts on making key attacks as difficult and expensive as possibleâ; deployed in a âproprietary setting, our main goal is to make black-box attacks computationally infeasibleâ; a âdetermined white-box adversaryâ is out of scope
- âBuilding (adversarially) robust watermarking systems remains an extremely challenging problemâ; âtraining a perfectly robust and secure watermarking scheme may be infeasibleâ
- âEventually there will be multiple versions in production ⊠vulnerability might be âinheritedâ between versionsâ (Sec. 7)
- Sec. 10: âSynthID-Image alone will not solve many of the problems we set out to alleviate, including misinformation, impersonation or copyright tracking ⊠watermarking in itself does not solve the provenance problemâ; wants âpublic detectability using cryptographic signaturesâ and better âsecurity, particularly considering white-box threat modelsâ for open models
what is deployed
- Google: SynthID, Google DeepMind
- âThe watermarks are embedded across Googleâs generative AI consumer productsâ; images and video âdesigned to stand up to modifications like cropping, adding filters, changing frame rates, or lossy compressionâ; audio from Lyria and NotebookLM âcanât be altered by common modifications like adding noise, MP3 compression, or changing the speed of the trackâ
- verification for the public is through chat: âupload the image, video or audio clip to your chat, and ask if itâs been created or altered by Google AIâ
- SynthID Detector, Google, May 2025: a portal âto early testersâ, waitlist for âJournalists, media professionals and researchersâ; âOver 10 billion pieces of content have already been watermarkedâ; NVIDIA Cosmos videos carry SynthID; GetReal Security can detect it
- the paper: âits corresponding verification service is available to trusted testersâ; the SynthID-O model is âavailable through partnershipsâ; nothing is open source except the text watermark
- OpenAI adopted SynthID rather than building its own
- Verify OpenAI-generated content, OpenAI, first archived 19 May 2026: âIt looks for supported signals, including C2PA metadata and SynthID watermarksâ; images and audio; âDetected signals are reliable, and false positives are rareâ; a âContent Provenance APIâ for developers; no-signal cases include âits watermark was degraded, it came from a legacy generation modelâ
- C2PA and SynthID in ChatGPT images, OpenAI help center, 2026 capture: âSupported images generated with ChatGPT, Codex, and the OpenAI API include both signalsâ; âSupported OpenAI-generated audio include an inaudible watermarkâ; metadata âcan sometimes be removed by platforms, editing tools, or file conversions. Watermarks ⊠may be more durable through some transformations. However, watermarks generally provide less context than metadataâ
- as of the SynthID-Image paper (Oct 2025) OpenAIâs promised DALL-E 3 watermark âhas not been releasedâ, so this is a 2026 change
- Sora: Launching Sora responsibly, OpenAI, Sep 2025: âall outputs carry a visible watermark. All Sora videos also embed C2PA metadata ⊠and we maintain internal reverse-image and audio search tools that can trace videos back to Soraâ; no invisible watermark claimed
- the visible logo is a measurement confound: RobustSora, Zhuo Wang, Xiliang Liu, Ligang Sun, arXiv Dec 2025, shows passive AI-video detectors partly learn the logo (âSora 2 induces drops of -11 to -14 ppâ when it is removed)
- Meta: Labeling AI-Generated Images on Facebook, Instagram and Threads, Meta, Feb 2024
- Meta AI images get âvisible markers ⊠and both invisible watermarks and metadataâ; cross-company labels use âthe âAI generatedâ information in the C2PA and IPTC technical standardsâ; âthere are ways that people can strip out invisible markersâ
- the SynthID-Image paperâs view in Oct 2025: Meta has âsome open-source models ⊠but no concrete public-facing verification system yetâ; Microsoft âopen-sourced an image watermarking modelâ (InvisMark)
- Adobe: TrustMark plus fingerprints plus C2PA as âdurable Content Credentialsâ (page above); the C2PA spec âspecifies measures to make the metadata durable, or able to persist in the face of screenshots and rebroadcast attacksâ
- laws pushing this
- EU AI Act Article 50, âComes into force 2 August 2026â: providers âshall ensure that the outputs of the AI system are marked in a machine-readable format and detectable as artificially generated or manipulated ⊠effective, interoperable, robust and reliable as far as this is technically feasibleâ
- California SB 942, bill text (via Wayback): âoperative on January 1, 2026â; covered providers (over 1,000,000 monthly users) must âmake available an AI detection tool at no cost to the userâ and offer manifest and latent disclosures
- China, Measures for Labeling AI-Generated Synthetic Content, translation by China Law Translate, in force 1 Sep 2025: Article 5 requires âimplicit labelsâ in âfile metadataâ, and âService providers are encouraged to add implicit labels to generated synthetic content in forms such as digital watermarksâ; Article 6 makes platforms check metadata and label content, which is the only law I saw that puts duties on the platform side
- Watermarks Without Verification: AI Text Watermarking After the EU AI Act, Alexander Nemecek, Vipin Chaudhary, Erman Ayday, arXiv Sep 2026, argues the real failure is that âno public tool can test the deployed systemsâ; the same holds for images
has anyone measured watermarked media in the wild?
Short answer: no, for invisible watermarks. The only in-the-wild numbers are about C2PA metadata and about passive detectors.
- I found no paper, report or blog that ran a watermark detector over a web crawl, a platform sample, or a news corpus; the reason is structural, every production detector is closed (Gemini chat, the Detector portal, OpenAI verify) and rate-limited
- GPT-Image-2 in the Wild: A Twitter Dataset of Self-Reported AI-Generated Images from the First Week of Deployment, Kidus Zewde, Simiao Ren, Xingyu Shen, Jiaqi Wu, Yuchen Zhou, Tommy Duong, Zikang Zhang, Ethan Traister, Kewen Xie, arXiv, Apr 2026
- 10,217 images from X in the week after the 21 Apr 2026 release
- âC2PA content credentials proved infeasible: Twitterâs CDN strips all embedded metadata on upload, leaving every image as a bare JPEG with no EXIF, XMP, or C2PA markersâ
- X shows a âMade with AIâ badge on â53.7%â of checked posts, so X reads provenance at upload and then throws the file-level signal away; most images are âserved at resampled resolutions by Twitterâs CDN, consistent with lossy transcoding on uploadâ
- they did not test SynthID, even though OpenAIâs images carried it by then; this dataset is a ready-made testbed for that
- passive-detector measurements exist and show the scale of what watermarks are supposed to cover
- AI-Generated Faces in the Real World: A Large-Scale Case Study of Twitter Profile Images, Jonas Ricker, Dennis Assenmacher, Thorsten Holz, Asja Fischer, Erwin Quiring, RAID 2024: ânearly 15 million Twitter profile pictures shows that 0.052% were artificially generatedâ
- Synthetic Politics, Zhiyi Chen, Jinyi Ye, Beverlyn Tsai, Emilio Ferrara, Luca Luceri, ACM Hypertext 2025: in 2024 US election tweets âapproximately 12% of shared images are detected as AI-generatedâ
- in my opinion, a platform study like the X one, repeated across platforms with watermarked originals we generate ourselves, is the cheapest high-value paper in this area; see ideas below
watermark vs C2PA, and whether a watermark is evidence
- Authenticated Contradictions from Desynchronized Provenance and Watermarking, Alexander Nemecek, Hengzhi He, Guang Cheng, Erman Ayday, CVPR 2026 Workshop APAI
- âa digital asset carries a cryptographically valid C2PA manifest asserting human authorship while its pixels simultaneously carry a watermark identifying it as AI-generated, with both signals passing their respective verification checks in isolationâ; needs âno cryptographic compromise, only the semantic omission of a single assertion field permitted by the current C2PA specificationâ
- lab study on 3,500 images; fix is a joint check
- AI Watermark Evidence Fails Forensic Readiness: An Empirical Evaluation, Saifur Rahman Tamim, Amir Labib Khan, arXiv Jul 2026 (text watermarks, but the framing transfers)
- laws assume âwatermark detection yields evidence reliable enough for courtsâ; SB 942 wants disclosure âpermanent or extraordinarily difficult to removeâ; after paraphrase âSynthID fared only slightly better at 98.3%â removal; âNone of the three methods satisfy more than two of five Daubert factorsâ
- Are Watermarks Bugs for Deepfake Detectors?, Xiaoshuai Wu, Xin Liao, Bo Ou, Yuling Liu, Zheng Qin, IJCAI 2024: watermark perturbations âare prone to overlap with the forgery signals used for detectionâ, so marking an image can change what a passive detector says
- Googleâs paper agrees that watermarks need company: âwe expect watermarking to be deployed alongside a metadata-based standard like C2PAâ and fingerprinting, because âsimilarity search is more prone to false positives rather than false negativesâ while watermarks fail the other way
what remains open
- adversarial robustness: no scheme survives a motivated attacker; the deployed answer is secrecy plus fingerprint databases, which only the company can query
- forgery: semantic marks are copyable; post-hoc marks are forgeable with detector access; defenses are weeks old
- public verifiability: SoK Zhao et al. ask âwhether schemes with public attribution and strong robustness can be efficiently instantiatedâ; Google wants âpublic detectability using cryptographic signaturesâ; nothing deployed does it
- open-weight models: a watermark in an open modelâs decoder is removed by anyone who can fine-tune; Google: âFor open models, we need to work on improving security, particularly considering white-box threat modelsâ
- versioning and interoperability: Article 50 says âinteroperableâ, but each vendorâs detector reads only its own mark (OpenAIâs page: content from âanother companyâs model ⊠the tool currently does not detectâ); OpenAI using SynthID is the first cross-vendor case
- partial content: âdetection and handling fractional watermarksâ (SynthID-Image Sec. 10) when a marked image is pasted into a collage or a video frame is cropped
- real-world survival: no public data on screenshots, messaging apps (WhatsApp, Telegram, WeChat recompress hard), platform CDNs, or print; all benchmarks simulate
- measurement: nobody knows what fraction of AI images on the web carry a working mark, or how fast marks decay as content gets reshared; and nobody outside the vendors can find out
- fairness: AudioMarkBench found ârobustness gaps among biological sex groups ⊠and language groupsâ; untested for images across content types beyond Googleâs âgrayscale photographs, sketchesâ note
research we could do
- each idea names the gap, what it builds on, the method, and the main risk
- the first three fit our web-measurement background and need no vendor access beyond public verify endpoints
- do watermarks survive the real web? a platform-pipeline study
- gap: every robustness number is simulated; the one real-platform observation (X strips C2PA) came as a side note
- builds on: GPT-Image-2 Twitter dataset method; SynthID-Imageâs transformation list; WAVES protocol; Adobeâs durability claims
- method: generate images and audio with Gemini, ChatGPT, Meta AI (all now stamp output); post them through X, Facebook, Instagram, Reddit, TikTok, WhatsApp, Telegram, WeChat, Weibo, email, Slack; re-download; also take phone and browser screenshots; check each result with the public verifiers (Gemini chat, OpenAI verify and its API) and with open detectors (Video Seal, WAM, AudioSeal) for our own re-stamped copies; report survival per platform, per transform chain, over reshare depth
- also check whether each platform strips, keeps or rewrites C2PA, which extends our C2PA work
- risk: closed verifiers are rate-limited and may change; mitigate by keeping the sample small (hundreds), logging verifier versions, and using open models for the dense sweep
- how much of the AI imagery on the web is watermarked, and what fraction is still readable
- gap: zero in-the-wild numbers; DeGenTWeb measures LLM text share of the web, this is the image counterpart
- builds on: DeGenTWeb crawl infrastructure; Ricker et al. and Chen et al. passive-detector pipelines; OpenAI Content Provenance API; Googleâs Detector portal (apply as researchers)
- method: sample images from Common Crawl or our crawl, from news sites, and from platform feeds; first pass with a passive AI-image detector to find candidates; second pass with every verifier we can reach; estimate the share of AI images that carry (a) C2PA, (b) a readable watermark, (c) nothing; stratify by site type and by time since Article 50
- risk: verifier access; false negatives of passive detectors bias the denominator; report bounds rather than point estimates
- the Article 50 natural experiment
- gap: nobody has checked whether the EU obligation changed what vendors and platforms actually ship
- builds on: âWatermarks Without Verificationâ argument; our labeling-rules note; idea 2âs pipeline
- method: snapshot the same generators and platforms before and after 2 Aug 2026 (the Wayback Machine helps for pages, but for media we must generate and test ourselves); record which vendors stamp, which platforms keep C2PA, which detectors exist and their terms
- risk: the âbeforeâ is gone for media; the study becomes a longitudinal one starting now
- a public, versioned watermark decay benchmark built from the GPT-Image-2 X dataset and its successors
- gap: benchmarks use synthetic distortions; real reshared copies of known-watermarked images now exist publicly
- builds on: Zewde et al. dataset (10k OpenAI images, which carried SynthID since 2026); OpenAI verify API
- method: for each image, query OpenAI verify; follow reposts and quote-tweets; measure how detection decays with each hop; release the per-hop results as a benchmark
- risk: OpenAIâs terms on bulk verification; X API cost
- forgery in practice: can a real photo be made to âverifyâ as AI at the public endpoints?
- gap: MĂŒller et al. and Saberi et al. show forgery against lab detectors; nobody tried against Gemini or OpenAI verify, whose output is a single bit and whose models are secret
- builds on: Saberiâs black-box spoofing (watermarked noise added to real images); re-watermarking paperâs observation that generic marks interfere; MarkNullâs claim against SynthID
- method: take SynthID-marked outputs, extract a transferable residual (average of watermarked minus regenerated pairs), add it to real photos, test at the endpoints with few queries; measure the false-positive rate achievable and the visual cost
- risk: ethics and terms of service; coordinate disclosure with Google and OpenAI; keep query counts low
- a joint C2PA plus watermark verifier for crawls
- gap: Nemecek et al. show the two layers contradict; nobody has built the joint check into a measurement pipeline or measured how often contradictions occur in real content
- builds on: Authenticated Contradictions protocol; our C2PA tooling; idea 2âs data
- method: implement the cross-layer audit; run it on crawled media; report the conflict matrix in the wild
- risk: needs watermark readers, so it inherits the access problem; can start with open marks (Metaâs) and C2PA-only states
- what do messaging apps do to media? a transformation atlas
- gap: SynthID-Image lists 30 transforms âreadily available on personal computers or smartphonesâ but nobody has characterized the actual transform chains of popular apps (resize targets, JPEG quality, chroma subsampling, video re-encode settings, audio codecs)
- builds on: SynthID-Image Sec. 4; Video Sealâs codec-in-the-loop training; our measurement habits
- method: send calibrated test media through each app and platform, infer the transform parameters from the output, publish the atlas; then plug the measured chains into WAVES, AudioMarkBench and VideoMarkBench so benchmarks test real chains
- risk: low; mostly engineering; apps change, so the atlas needs dating
- fractional and collage content
- gap: SynthID-Image names âfractional watermarksâ as open; WAM handles 10% regions in the lab; in the wild, AI images appear as thumbnails, memes with captions, and stills inside videos
- builds on: Watermark Anything; VideoMarkBench aggregation strategies; idea 1âs platform data
- method: build a test set of real collage, meme and screen-recording layouts from platform samples; measure detection as a function of marked-area fraction and scale; propose aggregation rules
- risk: again needs detectors; open models first
gaps in this review
- I could not read full texts for most 2026 attack papers (MarkNull, re-watermarking, speech-enhancement), only abstracts and arXiv HTML greps; numbers quoted come from abstracts unless a section is cited
- I did not find Amazonâs Titan image watermark documentation or any Midjourney, Stability or ByteDance statement; those vendorsâ deployment status is unknown to me
- Googleâs Gemini image-verification product pages returned 404 or JS-only, so the âask Geminiâ quote comes from the SynthID landing page only
consultation
- see the provenance index for the current Extra High consultation
Last edited: