logoalt Hacker News

stri8tedtoday at 5:39 PM2 repliesview on HN

This model was likely trained months before deepseek released their paper.


Replies

dannywtoday at 10:27 PM

Models are generally posttrained to a window shorter than you think.

manquertoday at 6:30 PM

Doesn't mean they didn't apply something similar. They could have also come up independently with their own version, the speculation is not they copied it, rather that they have performance breakthroughs which perhaps is a result of work in same domain