What % of time, for a an average session, do you think is app overhead vs waiting for tokens? And there's your answer for why it's not a priority.
From OpenAIs perspective, resources on your computer are free and wasting them is inconsequential
[dead]
Jon Blow's response to this take was, "yes, which is why you have to work even harder to hide latency", instead of adding more on top.