logoalt Hacker News

Retric • today at 2:24 AM • 0 replies • view on HN

> Yes at constant resources they are significantly better.

There has been some fundamental advancements in LLM training and efficiency. But good luck actually finding non toy models of exactly the same size, hardware, and training one from now and other from 2001 where the newer model is dramatically better at literally everything.

Most of the real world advances are from throwing ever more resources at the problem.

> Gen 2 takes

Diminishing returns are not a question of a single generation. Gen 2, 3, 4, 5… would also need to have the same 3x return on 2x resources or you don’t have an exponential curve.