logoalt Hacker News

neosatyesterday at 7:51 PM1 replyview on HN

Do you find the video understanding work there also to be 'silly little slop', or did you only look at the gifs on the page and not read about the understanding work in a 3B model?

This is not ground-breaking by any means, but achieving this in a 3B model and sharing the approach + weights advances engineering and certainly more contribution that 'silly little slop videos' imo.


Replies

MattRixyesterday at 8:08 PM

It’s not a 3B model, it has 3B active parameters. The full model is much larger.

show 1 reply