Is it? I think waferscale might actually be cheaper per-token, it's just so many more tokens, and of course right now it's not a full buildout so the availability is limited as well. I'd imagine they'll be migrating to whichever inference method is least expensive, and I expect asics to be the ultimate answer.
I'm not familiar with economics of chips, but I presume the SRAM on the wafer is less dense than HBM so it might be eating into its cost efficiency?