I always got the impression that C2PA is a way to say "this photo came from the BBC (for example) and they've only signed it because they've verified the supplied edit chain". It's always been obvious that one could point a camera at a screen, I don't think anyone involved with C2PA has claimed otherwise.
It seems like there's a big disconnect between what C2PA says it's for and what certain journalists think it's for.
Why not just have BBC sign stuff with their own creds then? You can originate the chain of custody anywhere and relying on the camera is kind of silly when the reputation of the original publisher in a more meaningful backstop.
C2PA is a bit nebulous as a concept, which is part of the problem. It is both of these things. The BBC type use case where a publisher signs their own content with their own keys seems reasonable to me, or at least, not obviously broken.
However I question the value-add when e.g. the BBC website is already authenticated by nature of being served over HTTPS, and anyone who redistributes BBC content can and should link back to the source.
> It's always been obvious that one could point a camera at a screen, I don't think anyone involved with C2PA has claimed otherwise.
They haven't claimed otherwise exactly, but some have implied it's a solvable problem. Here's where the "learn more" link goes, for when Youtube annotates a video as having C2PA metadata: https://support.google.com/youtube/answer/15446725 (Google is a C2PA Steering Committee member)
> The metadata that leads to a 'Captured with a camera' disclosure is made by a third party (for example, a camera manufacturer). This means that there is some risk that someone could take a photo of another screen showing synthetic content. Because the other screen shows an image that has been modified, it wouldn't be eligible for the 'Captured with a camera' disclosure. This issue is called 'air-gapping'. Camera manufacturers will continue to develop detection measures to prevent 'air-gapping', but the sophistication of those detection measures may vary in the near term.
Interestingly they do not mention any of the other known limitations. Their phrasing is highly weasel-wordy, but the implication is clearly that they imagine picture-of-screen detection to become robust (somehow) in the medium-to-long term.