I wish jev took in images so we could do this generically for any game, without memhacks. I'm sure that's coming.
You could front this with an image -> text model but that would be much lower quality vs latency, and the whole point of doing it with a decision model is remove the latency.
Games are a really interesting testing ground for robotics; if we can solve game playing (incl 3d) we could embody "system one" intelligence into robots that have something emulating general reflexes without needing to fine tune.
Agreed this would be super cool and I do see that coming in the future. But Jev-level latency just isn't there yet with full images.