logoalt Hacker News

omneity • yesterday at 2:55 PM • 2 replies • view on HN

I’m working on this problem using a vocab-free, byte-based approach. It’s definitely solvable.

https://huggingface.co/posts/omarkamali/593639295164067

https://huggingface.co/blog/omarkamali/tokenization


Replies

mdp2021 • yesterday at 9:16 PM

Careful: the problem is very certainly ___not___ counting letters. That is only a telling way to check "is the NN checking or not?". We demand that NNs for consultancy tasks check, strictly.

lern_too_spel • yesterday at 5:04 PM

I used to think byte level tokenization was the answer, but humans also think at a word level and only reevaluate the words at a character level when asked. The solution to better tokenization across languages is likely to be learned tokenization. Here is one attempt I have seen: https://github.com/SamD770/bitter-lesson-tokenization