ARC-AGI is a terrible benchmark for testing LLMs because LLMs are not made, trained, or tuned for playing games.
They are trained on text to respond well to text based questions and do tasks involving modifying text files.
They are not designed for playing games, looking at games, or visual puzzles. Also translating games into text input for the LLM skews the test completely.
Imagine trying to get a human to solve visual puzzle but they can’t look at the puzzle but it has to be explained to them in textual format, we would be terrible at it.
But yet we persist in wasting time on this benchmark. It doesn’t mean anything.