logoalt Hacker News

PunchTornadoyesterday at 7:43 PM1 replyview on HN

Look at those graphs another time. Claude beats gpt.


Replies

superfrankyesterday at 10:10 PM

Can you explain where you're seeing that? From what I see, the first two graphs have OpenAI models above Claude models (including Mythos) on the Technical Non-Expert and the Practitioner evals. Mythos now beats Codex 5.3 on the Expert eval and Opus was already on top for the Apprentice one although now Mythos leads there.

So, even including Mythos, OpenAI still has 2 models on top for the 4 evals listed.

show 2 replies