If this is true, it's even more impressive that some of the open weight models that are <3.5T in size, approx 33% of its size, are within a few points of it in the artificial analysis leaderboard.
And GLM is only 0.7T!
But these labs distill off the larger models. Both officially at the labs with the big ones, and unofficially. We need the giant models to get the smaller models.
You want to take a look at the "Scaling Laws" paper, so you can extrapolate from these numbers.
GLM 5.3 is "only" 753B parameters. Much much smaller.
Not necessarily, there could be diminishing returns on mere parameters count .
There is nothing to say for example a 1 Quadrillion parameter model will be vastly more intelligent than current SOTA especially since new training data is largely synthetic today