So, like ARC-AGI-3? I'm sympathetic to the notion that the things we're able to measure are necessarily going to miss important aspects of capability and intelligence. But people are attempting to measure this kind of thing, and models keep getting better at it.
And yet, curiously, they fall flat on their silicone asses until they have a copy of the test to "train on".