Moving from either frontier intelligence or frontier latency to a single model that does both at the same time is potentially a game changer in certain industries. I can easily see e.g. hedge funds dropping tons of money on this, because it means they can now do the same thing as their competitors, but much faster. That's basically a license to print money.
I am not sure this is the way to make AI more cost effective for such customers. If they are able to tweak any model for their use case it would be way more reliable and also way cheaper. In my opinion generic LLMs in the future will be just for attention economy or maybe government contracts. Everyone else will be running fine tuned free weight models or licenced closed source models (self hosted or managed).
What would a hedge fund want to do on this exactly? It’s too slow for hft and I’m not sure what they would be doing where ms matter but is not hft.