logoalt Hacker News

JonChesterfield • today at 11:12 AM • 2 replies • view on HN

Weird paper. Models have had tokens for each byte for ages now. They can read and write individual bytes just fine, in addition to also having multibyte tokens.


Replies

Philpax • today at 12:16 PM

The point is not to have byte tokens: it's to have only byte tokens, so that the usual failures of tokenisation (e.g. the number of Rs in strawberry) can be avoided.

➕ show 1 reply
andai • today at 2:18 PM

How can an LLM read individual bytes if all it gets is the tokenized ones? Or does it just get the ones that we didn't know how to tokenize?

➕ show 2 replies