logoalt Hacker News

Retrofitting language models to operate over bytes

103 points • by theanonymousone • last Thursday at 2:15 AM • 23 comments • view on HN

Comments

armcat • today at 9:59 AM

There has been extensive research into token-free LLMs, but for some reason we are still operating in a token domain, so there is something to that.

Byte Latent Transformers (BLT): https://arxiv.org/abs/2412.09871

Charformer: https://arxiv.org/abs/2106.12672

➕ show 3 replies
mrkn1 • today at 9:34 AM

subword tokenizers were never causal in the first place, so BPE was peeking at future bytes all along! TIL

puttycat • today at 2:35 PM

I don't understand how this is different from BPE tokenization.

serioussecurity • today at 4:47 AM

Wow nature got rolled. Should have stayed closer to their expertise. They were already being hustled by a lot of the applied AI work they were accepting.

➕ show 1 reply
BitProgram • today at 8:04 AM

[flagged]

singularityisne • today at 6:02 AM

[flagged]

JonChesterfield • today at 11:12 AM

Weird paper. Models have had tokens for each byte for ages now. They can read and write individual bytes just fine, in addition to also having multibyte tokens.

➕ show 2 replies