logoalt Hacker News

fastballyesterday at 8:18 PM1 replyview on HN

Tokenization is <0.1% of the inference time for the first token in the same way it is <0.1% for the last.


Replies

marcelroedyesterday at 8:57 PM

Time to first token refers to the time until the model outputs one token, which includes the time to process the entire prompt (doing prefill). The GPU time per token is much lower when doing prefill, so the significance of tokenization is higher.

show 1 reply