Beware of the benchmarks listed. SciCode and EnterpriseOps for instance: https://shukla.io/blog/2026-08/gym.html
The Chinese models also like to cut corners on stuff like science. Their scores on stuff like biotech and scientific knowledge is far from ChatGPT unfortunately. (Claude is pretty good but it just refuses all prompts).
The Chinese models also like to cut corners on stuff like science. Their scores on stuff like biotech and scientific knowledge is far from ChatGPT unfortunately. (Claude is pretty good but it just refuses all prompts).