Crazy it's still the only video understanding endpoint. It's what I use it for and no other model even offers a competitor.
You probably can't build such a model without unlimited access to YouTube and Google has been tightening the screws on that over the years pretty systematically.
to be fair, all it's doing is sampling the frames and maybe doing transcription, if I'm not mistaken. So you can do it with the other models too, you just need to sample the frames yourself and do the transcript yourself