Why would they need help figuring that out? I can fully believe a decent LLM would figure this out on its own.
I had a flash model without vision capabilities take screenshots and convert them to ascii to "see" what was going on, all on its own. That's just one example. They're very determined.