logoalt Hacker News

BinRooyesterday at 11:21 PM1 replyview on HN

Beware of the benchmarks listed. SciCode and EnterpriseOps for instance: https://shukla.io/blog/2026-08/gym.html


Replies

Onavoyesterday at 11:49 PM

The Chinese models also like to cut corners on stuff like science. Their scores on stuff like biotech and scientific knowledge is far from ChatGPT unfortunately. (Claude is pretty good but it just refuses all prompts).