When the market settles down, most of the money is going to be in bulk inference, for batch jobs. In 5 years everyone's laptop will be able to run Qwen 3.8 27B for coding tasks, but businesses will still need to run inference 24/7. Very few people will still need SOTA models once you can run an Opus 4.6 class model on your laptop.