logoalt Hacker News

theszyesterday at 3:45 PM0 repliesview on HN

Byte Pair Encoding [1] will be different for different languages. Application of the per-language BPEs to the input text will produce encodings with different lengths.

[1] https://en.wikipedia.org/wiki/Byte-pair_encoding

It naturally takes care of common prefixes and suffixes.

It is easy and fast to apply using radix tree or with finite automata. Even without radix tree, it is possible to have processing speed in the range of hundredths of thousands of bytes per second.