logoalt Hacker News

zmmmmmyesterday at 10:23 PM1 replyview on HN

it's great but we need a multi-modal model of this quality and price to truly declare victory.

But it makes me quite curious, how a text-only model can do so well on ARC-AGI-2 being a set of visual puzzles? It would have to solve it entirely using text-only spatial reasoning about the grid (or maybe writing code?). I am curious if this is normal or do other models use their vision capabilities to solve the puzzles?


Replies

ghosty141yesterday at 10:46 PM

xiaomi mimo is very cheap and not bad.