Another interesting discrepancy is that people think current GPUs are maybe capable of running AGI but they still can barely manage photorealistic rendering of a single room in realtime, or simulate something like a shirt thrown into a pile of laundry. They can generate a video of it based on millions of existing videos, but not do a real simulation of light and physics in realtime.
An AGI doesn't need that level of detail to do most tasks effectively. To use the OP's example, a cat catching a bug out of the air does not need to run a fluid dynamics simulation of airflow over the bug's wings to be able to catch it. A cheap approximation of the flight path is sufficient. Perhaps some physical tasks will need that level of detail but many will not.
Is the human brain capable of doing real simulations of light and physics in realtime? Or does it hallucinate the details and represent some low-resolution mental image? You may be overestimating the capabilities of flesh-based neural networks and underestimating the capabilities of silicon-based neural networks.