From blog post:
“It generates both image frames from novel views and explicit 3D outputs”
Model can indeed generate novel views but I don’t think in real time (I could be wrong). If you want to navigate a space in real time as user above was asking gotta rely on the 3D output and the explicit representation will provide continuity. User likely was referring to the limitations we’ve seen on video-style world models like genie and precursors.