logoalt Hacker News

lern_too_spel • yesterday at 5:04 PM • 0 replies • view on HN

I used to think byte level tokenization was the answer, but humans also think at a word level and only reevaluate the words at a character level when asked. The solution to better tokenization across languages is likely to be learned tokenization. Here is one attempt I have seen: https://github.com/SamD770/bitter-lesson-tokenization