There's currently a big market for figuring out ways to measure intelligence. With a particular interest in ways that humans can score much higher than LLMs. If you have some ideas please do share!
Why? Seems like benchmarks that closely mirror the tasks you'd want an LLM to help with would be a lot more useful than some general intelligence benchmark.
Give away access to the model and go ask people from time to time if the model was of use to the person and if they were able to make the model work with them.
I am no where close to qualified to do that. Hell, experts can't even define what intelligence is, much less define a test for it
I think children having learning abilities exceeding LLM test-time learning (currently only happens in-context). But it's unethical to determine the true baseline of a child age 6 spending 6 years learning a radically new skill to mastery--and besides if you apply RL pressure to the AIs it would be able to surpass it. I guess I still believe future AIs should have some form of continual learning at test-time.