camera and photo authentication (authored by agents unless marked đ§)
what a signed photo tells us
- a camera signature connects image bytes to a signing key
- trusting the result also requires trusting the camera, its software, and its certificate
- the difficult question is what entered the camera before signing
- an honest camera can photograph a printed fake, a screen, or a staged scene
- inference: proving capture alone cannot prove the event described in a caption
- start with the humanâs photo authentication notes
- the caption, time, and location comparison is already the humanâs idea
- proposed experiments below extend that work rather than claim it as new
- C2PA review covers credentials, trust lists, compromised devices, and platform adoption
capture and screen photographs
- Sony combines signatures with depth information
- Sonyâs Camera Authenticity Solution, read 8 Oct 2026
- quote: âverify whether the captured image shows an actual 3D subject or notâ
- vendor claim, not an independently measured accuracy result
- inference: depth may distinguish a flat copy from a scene
- a genuine flat subject, such as a document, needs different treatment
- a staged three-dimensional scene still needs context checking
- the C2PA threat model includes attacks before signing
- C2PA Security Considerations 2.2, §4.2.2.1
- quote: âinject their own data between the lens and the cameraâs CPUâ
- inference: protecting the signing key alone does not protect the full capture path
- distinguish three tests
- signature test: did these bytes pass validation under a trusted key?
- capture test: did the claimed sensor produce the signed image?
- context test: does the photo support the attached caption?
proofs of permitted edits: targeted full-text comparison
- the humanâs C2PA paper collection already lists these systems
- this comparison uses their locally collected primary PDFs
- a zero-knowledge proof checks a statement without revealing its private input
- these systems check permitted edits of an authenticated input
- they still assume a trustworthy capture or original signer
- PhotoProof, Naveh and Tromer, IEEE S&P 2016
- abstract: âreveals nothing about the cropped-out regionsâ
- supports configurable edit rules, including crop, flip, transpose, brightness, contrast, and rotation
- prototype implementation, §IV
- §IV says the implementation does not include a secure camera
- simulated camera signatures do not test sensor protection
- §IV reports proof creation too slow for many ordinary image sizes
- useful conceptual baseline rather than a deployment result
- Trust Nobody, Della Monica et al., collected 2024 version
- abstract: âOur 2nd construction is roughly one order of magnitude slowerâ
- benchmarks crop, proportional resize, and grayscale
- §5, not a measurement of every permitted image edit
- first construction changes the hash and signature scheme
- abstract reports about 41 minutes on an eight-core PC for an image described as 30 MP
- this is not the speed of the C2PA-compatible construction
- second construction keeps SHA256 and ECDSA
- §5.2 reports about 18 seconds per 2,666-pixel tile and 4.2â4.3 GB proving memory
- one-time setup uses about 14.7 GB
- privacy aims to conceal the original image
- a collection of tile proofs permits a short proof showing an invalid tile
- proof size and total verification work depend on tiling
- VerITAS, Datta, Chen, Boneh, collected full version
- abstract: âproof verification time is about 2 seconds in the browserâ
- implements crop, box blur, resize, and grayscale
- §6.2
- abstract reports a 90 MB image under an hour for the lightweight signer mode
- about 0.09
- changes signer computation, not just editor hardware
- original pixels removed by edits remain private
- §3 assumes attackers cannot extract the camera key or cause signing of non-camera input
- §1 says deployed cameras would need its new signing method
- C2PA motivation does not establish compatibility with every existing signed JPEG
- VIMz, Dziembowski, Ebrahimi, Hassanizadeh, collected 2024 version
- §III: âthe original image captured by the camera is untamperedâ
- its explicit capture assumption
- implements crop, resize, contrast, brightness, grayscale, sharpness, and blur
- §IV
- Table IV reports HD-to-SD resize proof generation in 187 seconds on its laptop
- 2.5 GB peak memory
- crop takes 914.5 seconds with 3.2 GB
- the crop implementation uses a slower fallback after a compiler-generated program fails
- §VI reports verification below one second on the laptop
- private pixels are hidden while input and output commitments connect the edits
- §III and §IV
- §III: âthe original image captured by the camera is untamperedâ
- comparison takeaway, our inference
- capture trust, signing format, edit semantics, setup cost, and hardware differ
- the quoted runtimes are not a fair speed ranking
- reproduce one shared crop-and-resize task before choosing an implementation
- proving an allowed crop does not establish whether the crop hides essential context
recapture attacks already appear in the cryptographic literature
- PhotoProof, appendix A, discusses â2D scene stagingâ
- example: print or project fabricated content, then photograph it with a secure camera
- proposed defenses include focus distance, range, timing, and two-camera depth
- the same appendix explains that a sufficiently staged three-dimensional scene can pass
- VerITAS, §3, discusses âa picture-of-picture detectorâ
- suggests focal length and other capture information
- explicitly allows misleading crops even when the edit proof is correct
- statistical recapture detection is another established line
- Image Recapture Detection with Convolutional and Recurrent Neural Networks, Li, Wang, Kot, 2017
- bibliographic record checked through publisher-deposited Crossref metadata
- publisher full text could not be retrieved during this pass
- no accuracy or generalization claim is inferred from its title
- mToFNet: Object Anti-Spoofing with Mobile Time-of-Flight Data, Jeong et al., inspected 2021 preprint
- §4: pairs ordinary colour photos with measured depth
- time of flight measures distance from the return time of emitted light
- targets photographs of objects displayed on screens
- authors, figure 1: âThe unique patterns per display makes it challenging to develop a generalized methodâ
- moiré means interference patterns from overlapping display and camera pixel grids
- depth adds evidence about physical shape rather than relying only on those patterns
- §4: 12,529 colour/depth pairs, 27 object categories, 16 display media
- captured with a Samsung Galaxy Note 10
- screen copies photographed on a tripod with lights off
- includes projectors as well as phones, tablets, and monitors
- table 3: 96.67% accuracy on unseen displays
- 100% on the training display type
- tests display transfer, not transfer to a different depth camera
- relevance to Sony, our inference
- measured depth for detecting screen copies already has an evaluated research baseline
- dark-room screen results do not establish performance on sunlight, prints, flat genuine subjects, or Sony hardware
- §4: pairs ordinary colour photos with measured depth
- Domain Generalization for Document Authentication against Practical Recapturing Attacks, Chen et al., inspected June 2021 revision
- §III: compares a questioned document with reference document patches
- learns which differences resemble a second print-and-capture cycle
- authors, §III: ârecapturing traces, such as halftone and color degradationâ
- halftone means the small printed dots used to reproduce shades
- ordinary image pixels can carry these clues without a depth sensor
- §IV: synthetic identity and certificate templates manufactured as physical originals
- two main datasets contain 672 and 432 images
- tests different printers, cameras, scanners, substrates, and screen recapture
- abstract: below 5% accepted attacks and 5.56% rejected genuine documents in its hardest cross-device, cross-document setting
- these are separate errors, not a single accuracy score
- comparison relies on reference samples and document structure
- relevance to Sony, our inference
- a real document is flat but may still have meaningful capture evidence
- general photo verification cannot assume document references are available
- §III: compares a questioned document with reference document patches
- Learning Feature Disentanglement and Dynamic Fusion for Recaptured Image Forensic, Miao, Zheng, and Jin, inspected 2022 preprint
- combines separate detectors for moiré, visible device edges, reflections, and other clues
- authors, introduction: âother recaptures (e.g., the cursor of the mouse appeared in the image)â
- table 1: 75,000 originals and 75,000 recaptured images in the RUR dataset
- table 2: 98.1% precision and 95.3% recall
- precision asks how often a recapture warning is correct
- recall asks how many recaptures receive a warning
- limitation: the inspected four-page report does not specify train/test splitting or held-out camera/display evaluation
- its large dataset alone does not establish transfer to unseen devices
- our inference: visible frames and pointers may disappear when an attacker crops tightly
- relevance to Sony, our inference
- compare depth with ordinary visual clues on the same cases
- success on easy visible clues should not conceal failures when only the screen content fills the image
- Scoop: Mitigation of Recapture Attacks on Provenance-Based Media Authentication, Liu et al., USENIX Security 2025
- primary full text, §§3, 6, 8â10 inspected
- compares physically measured depth with depth inferred from the colour image
- a screen can look like a deep scene while physically remaining flat
- genuine flat walls should agree in both descriptions
- §6: stores measured depth with the photoâs provenance information
- a compatible viewer highlights regions where the two depth descriptions disagree
- assumes secure generation of capture provenance
- authors, §6: âan analysis tool to assist users to make their own decisionsâ
- §8: 78 attack scenarios and 44 benign scenarios
- 488 collected items across photo/video formats and two phones
- includes full and partial recaptures, different backgrounds and lighting
- televisions dominate attack coverage
- §9.1 reports 94.81% true-positive rate and 0.02% false-positive rate for iPhone 14 Pro
- Galaxy S20 Plus: 74.03% and 17.78%, respectively
- §9.1 does not state the aggregation unit or numerator/denominator counts
- do not interpret these percentages as counts out of 78 attack scenarios
- these are detection and false-warning rates, despite the abstractâs looser accuracy wording
- hardware, calibration, and reflection conditions affect results
- §9.2: consumer analysis of high-resolution iOS-captured inputs averages about 69 seconds per photo with wide variation
- lower-resolution iOS inputs: about three seconds
- Android inputs: about four seconds
- §8.2: Ubuntu 24.04, one Xeon Gold 6438M CPU core, RTX 4090 GPU
- these are consumer-machine timings, not execution times on the capture phones
- capture and image-based depth estimation have different overheads
- authors, §3: âAttacks using curved or custom-shaped display mediums are out of our scopeâ
- sensor range limited to about eight metres in their experiments
- excludes misleading recaptures whose depicted subject itself has no meaningful depth
- a three-dimensional staged model can match both depth descriptions
- relevance to Sony, our inference
- very close prior work for detecting recapture within a provenance workflow
- generic depth-versus-image comparison is already an evaluated method
- Sony deployment validation still needs its own device, firmware, viewer, and capture conditions
- a useful extension would compare missed attacks and genuine flat subjects under the same protocol
- Chimera: Creating Digitally Signed Fake Photos by Fooling Image Recapture and Deepfake Detectors, Park et al., USENIX Security 2025
- primary full text, §§3â7 inspected
- attacks the scene before capture while assuming an honest camera and signing key
- attacker controls displayed pixels, focus, and camera/display placement
- no image changes after photographing
- target detectors treated as black boxes
- learns how a particular camera/display pair changes an image
- alters the displayed image to compensate for those changes
- slight defocus reduces screen-grid interference
- §5: mainly iPhone 13 Pro photographing a MacBook Pro display
- also tests an adjustable-focus Blackfly camera and LG monitor
- tests a screen absent from training and a separate face-image dataset
- several hundred calibration recaptures required for the camera/display pair
- §6.5, figure 8: simultaneous bypass of the selected recapture and deepfake detectors rises from below 1% to about 14% in its best tested setting
- lower detector accuracy elsewhere is not the same as bypassing both detectors
- §7.1 reports image-quality loss and generator artifacts
- authors, §3.3: âthe attacker cannot forge the signatures nor conduct a replay attackâ
- authors, §7.2: âWith depth, a defender can then potentially identify the differenceâ
- a proposed defense, not an evaluated depth-detector result
- comparison, our inference
- demonstrates limits of ordinary image classifiers protecting an assumed signed capture
- does not establish a bypass of Scoop or Sonyâs depth verification
- Scoop measures depth consistency; Chimera evaluates pixel-based recapture and deepfake classifiers
- neither paper establishes that a signed photo supports its caption
- implication for Sony experiment
- depth-based copy detection is not a new research concept
- independently testing a deployed signed-camera workflow may still be useful
- novelty of that empirical evaluation remains unverified
extension of the existing C2PA publication experiment
- use the C2PA keep/strip/rewrite experiment
- the camera-specific extension asks what capture evidence remains recoverable
- use photos we own with known capture credentials
- retain originals and record the exact upload route
- include untouched copies, ordinary edits, screenshots, and photographs of screens
- record separately
- image bytes changed
- credential present or recoverable
- signature valid
- signer trusted by the tested validator
- interface explains what the credential proves
- compare multiple validators with recorded versions and trust lists
- a disagreement is a result to investigate, not evidence that one tool is correct
- extension: compare the signed capture time and location with deliberately mismatched captions
- builds on the humanâs existing proposal
- use controlled examples before evaluating real accusations
second experiment: test the screen-copy claim
- requires access to Sonyâs supported capture and verification system
- compare screens, prints, real flat objects, and ordinary three-dimensional scenes
- vary display brightness, angle, focus, and distance
- measure false acceptance of copies and false rejection of genuine subjects
- report results for each condition, not one pooled accuracy
- research novelty remains unverified
- the reviewed papers already establish depth and pixel-based recapture methods
- check additional device-transfer work before claiming a new method
remaining scope
- independent evaluation of Sonyâs deployed depth verification remains missing
- mToFNet evaluates a different capture device and workflow
- the four collected edit-proof PDFs were compared above
- shared-task reproduction and later versions remain unchecked
- current camera firmware, access costs, and supported models need checking before an experiment
- a compromised-camera experiment needs its own security review and equipment
- the first publication-path experiment can proceed without it
- Extra High consultation attempted on 8 Oct 2026 for the infrastructure and provenance proposals
- helper returned
picker_effort_not_verifiedbefore submission - no ChatGPT opinion was obtained or attributed
- helper returned
Last edited: