logoalt Hacker News

yt1998 • today at 1:59 PM • 2 replies • view on HN

Why chose CLIP to do this. Have you tried small VLMs like Qwen-VL? I believe those models have video encoders can better perform at this scenario.


Replies

radicality • today at 5:24 PM

Not OP, but probably because they have no idea what they are doing or what CLIP even is. And the whole thing is likely just from one llm prompt, and then they dumped the whole thing onto GitHub, undoubtedly without even looking at any of the code or architecture.

And then you end up with image metadata processing code like this, which just by a cursory glance I’m sure has edge case bugs : https://github.com/allenv0/SCM/blob/main/screenshot-probe.js

allenleee • today at 5:08 PM

[flagged]