to be fair, all it's doing is sampling the frames and maybe doing transcription, if I'm not mistaken. So you can do it with the other models too, you just need to sample the frames yourself and do the transcript yourself