logoalt Hacker News

numpad0yesterday at 6:37 PM1 replyview on HN

Those labs publicly said during GPT-3/4 era that the optimal epoch count, or dataset repetition count, for foundation model training, is one. So it's a forward 1-pass compression.

But it's a black box! Nobody knows whats going on inside! It's all transformative! Sure...


Replies

paulddraperyesterday at 8:03 PM

Have they seen big pharma?