logoalt Hacker News

armcat • today at 9:59 AM • 3 replies • view on HN

There has been extensive research into token-free LLMs, but for some reason we are still operating in a token domain, so there is something to that.

Byte Latent Transformers (BLT): https://arxiv.org/abs/2412.09871

Charformer: https://arxiv.org/abs/2106.12672


Replies

razodactyl • today at 12:53 PM

It's because of "chunking" which human minds do as well. If you increase granularity, you increase permutations and required compute to match the same performance of existing systems.

GPT4's tokeniser for example made use of extended tokens for coding structures and with enough training on code was quite ahead of the rest for a while.

The tokeniser and its vocabulary makes a big difference.

➕ show 1 reply
gchamonlive • today at 11:50 AM

Aren't diffusion models token-free?

➕ show 1 reply