logoalt Hacker News

worldsaviorlast Sunday at 9:55 AM1 replyview on HN

That's a 2.4T model, how would they reduce this to 35B and still give some accuracy? That's a completely different arch.


Replies

cyanydeezlast Sunday at 11:24 AM

there's been a lot of research about reducing models by taking out layers; there's also using it to train smaller models by optimizing parameters.

I dont see most model building as anything more than a pig at a slop troth, despite the level of sophistication; they're still rarely pruning the input beyond random sampling.