All I’ve been able to do is mentally tag parts of a shape as being at different “depths” (possibly with two parts at the same 3D location but tagged as having different “depths”). Which, yeah, I agree isn’t really visualizing a 4D shape. I think someone better at it than me might be able to use this to leverage one’s visual spatial intuition in a way that is a bit closer to what 4D visualization might be like if it were possible.
But I don’t think it is possible for us. At least, for strict senses of visualization. I suspect that the architecture of our brains doesn’t support it. This is rather speculative because I don’t know basically any neuroscience, but I wonder if the whole “the brain is largely a very wrinkled surface (though with several layers to this surface), and the retinas are also surfaces” may have something to do with this. A retina in a eye-4-ball would have a 3D boundary (I.e. hypersurface), and I imagine that the amount of information that it would take in if it had an at all similar angular resolution, would be too great for our visual cortex to be able to represent all of it.
Of course, our vision only really has the high resolution it seems to everywhere in the narrow region we’re directly looking at, but still.