it's great but we need a multi-modal model of this quality and price to truly declare victory.
But it makes me quite curious, how a text-only model can do so well on ARC-AGI-2 being a set of visual puzzles? It would have to solve it entirely using text-only spatial reasoning about the grid (or maybe writing code?). I am curious if this is normal or do other models use their vision capabilities to solve the puzzles?
xiaomi mimo is very cheap and not bad.