To do this you need video/motion understanding, the intent cannot be judged from still images or state descriptions.
We’ve had the tool to do this since mid 2025, V-JEPA2 [1], Yann Lecun’s last work at Meta.
It runs at several FPS on a macbook and can even be trained locally. Chaining it with Jev for decision-making would probably work great!
The technique i mentioned in the blogpost works with video too. I tested succesfully with Qwen/Qwen3-VL-8B-Instruct. Effecient caching is a little trickier though, but very doable.