(March 2024) Antropic claimed Claude 3 Opus had "graduate-level expert reasoning" with GPQA results of around 60% showing a roughly phd level performance.
(Sept 2024) OpenAI claimed o1 was phd-level in their launch post.
You're kinda both wrong. :)
They claimed the model was PhD-level, but they never mentioned the university the model graduated from... :)
They claimed the model was PhD-level, but they never mentioned the university the model graduated from... :)