logoalt Hacker News

JaRailyesterday at 12:36 PM1 replyview on HN

(March 2024) Antropic claimed Claude 3 Opus had "graduate-level expert reasoning" with GPQA results of around 60% showing a roughly phd level performance.

(Sept 2024) OpenAI claimed o1 was phd-level in their launch post.

You're kinda both wrong. :)


Replies

rescbryesterday at 2:54 PM

They claimed the model was PhD-level, but they never mentioned the university the model graduated from... :)