What about do you mean by single threaded? Each token is predicted by using parallel computation on the GPU.
Multiple agents need tokens. Should optimize for that instead of one agent blocking the others.
Multiple agents need tokens. Should optimize for that instead of one agent blocking the others.