logoalt Hacker News

Exploring Claude/GPT Knowledge Cutoffs and Pre-Training Timelines

144 pointsby sshh12yesterday at 2:20 PM21 commentsview on HN

Comments

myworkaccount2yesterday at 6:06 PM

I wonder if this kind of analysis will give us a way to check if the frontier labs are waiting for the right moment to release their models. To me it feels obvious that these companies are not releasing models as soon as they are done doing their post training / testing with any new model.

But there is no real way to know how much of this "waiting" any lab is doing, if we can get better estimates this way maybe we can gauge how far the open weights models really are.

show 2 replies
rad-byesterday at 5:35 PM

Great read and interesting analysis! I’m less charitable toward Anthropic supposedly not distilling ChatGPT for training purposes. Maybe not today, but during the GPT-4 era when Anthropic was the underdog - I can see it happen. Packaged along with some of Amodei’s clever jumping through hoops to prove how that is, in fact, virtuous.

ddxvyesterday at 4:20 PM

This was great! One thing that wasn't addressed that I always assume, is that a marketing name like "Opus 5" is not a single model, but many models, versions and gets minor updates over time.

I also assumed many questions get routed to simpler models or programs to answer correctly, but it almost surprisingly didn't seem that way from the post.

Anyways, great post.

show 1 reply
tobwenyesterday at 8:07 PM

From my own experiments with various models, I suspect that LLMs have distinct/partitioned cutoff dates; for example, historical literature doesn't change (Greek history, Shakespeare, Goethe), general knowledge (updated only in certain areas), technologies (updated regularly), software also remains surprisingly stable - for example, with GIT, a basic command set is sufficient to do 99% of the jobs - new features are unknown or unnecessary, and tabloid knowledge, which is always up-to-date (politics, Taylor Swift albums).

show 1 reply
sashank_1509yesterday at 4:57 PM

So Jan 2026 latest, I guess we are due to 2 OOM’s better retraining over the coming years. Curious to see the point at which AI plateaus.

brcmthrowawaytoday at 4:43 AM

Why don't the labs just run autoresearch on improving themselves?

show 1 reply
fosterfriendsyesterday at 7:32 PM

great read

engzaaninyesterday at 8:56 PM

[flagged]