I think the most important insight is the limitation of LLM perception:Slowly taking screenshots.
That method of perception probably scales N^2... so sure with more compute, LoTR animation will improve. But I think to get a real jump in "experiential feedback", perception needs to scale linear or sublinear. Maybe that's there LeCunn's jepa will come in.
There needs to be the removal of the middle man:
image -> text -> action
To image -> action.