I stand corrected, only Google's post-Ironwood TPUs have the split as well.
Nonetheless, TPU architectures are still a systolic array, and have their own limitations for scalability and flexibility. CUDA is no silver bullet, but it satisfies the demands of the edge and research customers very well.