logoalt Hacker News

sosodevyesterday at 9:18 PM1 replyview on HN

What about do you mean by single threaded? Each token is predicted by using parallel computation on the GPU.


Replies

josh-wraleyesterday at 9:42 PM

Multiple agents need tokens. Should optimize for that instead of one agent blocking the others.

show 1 reply