Yes at constant resources they are significantly better.
Diminishing returns aren't necessarily a problem if the rate of increase in resources is faster. In other words, if Gen 2 takes twice the resources but Gen 1 figured out a way to triple compute efficiency then there is no ceiling.
> Yes at constant resources they are significantly better.
There has been some fundamental advancements in LLM training and efficiency. But good luck actually finding non toy models of exactly the same size, hardware, and training one from now and other from 2001 where the newer model is dramatically better at literally everything.
Most of the real world advances are from throwing ever more resources at the problem.
> Gen 2 takes
Diminishing returns are not a question of a single generation. Gen 2, 3, 4, 5… would also need to have the same 3x return on 2x resources or you don’t have an exponential curve.