> But people have been saying this about CUDA for twenty years, and we are not any closer to a replacement GPGPU paradigm today.
How much money was in it for the first decade or so? I think AMD was asleep at the switch but e.g. Apple just did their own thing for the parts which they prioritized.
My understanding is also that Anthropic and OpenAI have also worked to decouple themselves so I think it’s likely that the CUDA moat is going to be less of a barrier than it used to be from the perspective of guaranteeing Nvidia profits.
Money wasn't really the problem. Apple pulled OpenCL together pro-bono, and worked with Khronos to find willing industry stakeholders that would oppose Nvidia. OpenCL needed hardware standardization though, and nobody wanted to design or implement on a scalable GPGPU architecture like CUDA had. AMD and Apple both bet big on raster efficiency, which turned out to be a terrible play when Nvidia was already putting dedicated ray tracing and tensor hardware into their GPUs. They both bet the farm against each other, and only Nvidia won.
Once Apple fully left Khronos, AMD played the smartest card they had; they architecturally split RDNA and CDNA into separate product lines, so they could optimize them independently. This staunched the bleeding, and gave AMD a datacenter presence that Apple Silicon could only dream of. Still not a scalable architecture, but better than nothing.