logoalt Hacker News

gcryesterday at 11:38 AM1 replyview on HN

This is oversimplified. The proposed question is whether a model with fewer parameters could achieve performance on one language similar to that of a larger model that’s been trained more broadly, which isn’t straightforward to do.


Replies

msdzyesterday at 2:09 PM

Oh, I didn’t interpret the above question as asking in that direction; but yeah, that’s of course something I didn’t attempt to answer with my comment.

Although I’d be intrigued in the answer to that small-narrow vs. large-broad model question, too!