logoalt Hacker News

padolseyyesterday at 6:55 PM2 repliesview on HN

The smaller these frontier-nearing models get, the more I'm reminded of https://en.wikipedia.org/wiki/Lottery_ticket_hypothesis


Replies

keeganpoppenyesterday at 6:59 PM

i think there definitely is some truth to this in terms of embeddings spaces, which is why i believe they are implemented by OpenAI/Anthropic in roughly highest import => least import bit order-- an overwhelming majority of the variance is in the first few hundred vector bits. i haven't actually tested this myself by manually truncating vectors, but it is my understanding that they generally speaking have this property.

Havoctoday at 7:41 AM

I wonder whether it’s possible to test the seed on a small model and then size it up on whatever works