"AI" has limits in that it cannot invent knowledge, it can only distill and search for patterns in existing knowledge
not sure how many will get this reference but "AI" for science and math is like super-shoes for runners
at first we are blown away by the impossible improvements including sub-2-hour realworld marathon and every other PR/CR/WR is dialed down
but then the improvements slow and reach a stall point because of the limit of technology and the source of the achievement
ie. sub-2-hour marathon yes, sub-1-hour never happening (rollerblade inline-skate record is 1-hour marathon)
Not sure what your first sentence means, or why you are quoting AI. Many of these problems individually were math at a level approaching the highest possible for expert human mathematicians. These are not simple combinations of existing ideas, or following of human intuitions, or implementing something following specific human instructions. Then again maybe you mean that math is not knowledge and that all math is simply extending the basic axioms using known patterns, to which I would not agree.
as much as I love analogy with humans using super-shoes, not everybody can be a world-class expert in their industry. There should be a place at the table for average people to take part; otherwise, it won't be sustainable.
As a programmer, I am mostly interested in whether my role is sustainable long-term and whether the models will get better. I don't feel in jeopardy yet, but two more years like this and the calculus of hiring software engineers could shift even further. QAs are already overwhelmed with work
How do humans "invent knowledge"? Is your argument that the answer to these questions already existed in the training set? Why didn't any human recognize that before?
I agree. The most likely scenario is that this is just a "new normal" lift that is percolating through human endeavors and will saturate at some point. For example, the whole cyber-security bruhaha should ultimately resolve into higher standards for code published -- we can now cheaply find and fix a whole slew of minor bugs that weren't worth our time before.
The fact we see a lift is not the same as evidence that the lift is unbounded.
The lift being finite is supported by the fact improvements have come at the edges: improvements from human feedback, improvements in harnesses, improvements on model compatibility with harnesses, improvements in inference efficiency with new architectures, etc. If we were just training better models from scratch that would be one thing, but we are just making better use of a tool we've developed.