logoalt Hacker News

stingraycharlestoday at 10:35 AM2 repliesview on HN

Back in the day - maybe two decades ago - I implemented language detection like this.

I seeded gzip compressors’ dictionaries with Wikipedia articles in different languages.

I would then try to use said dictionaries on any random text, and the one that was best able to compress it, was the correct language.

Absolutely totally not the best approach, but very fast and super simple to implement.


Replies

ape4today at 2:02 PM

Or maybe make a list of the most used 1000 words in each language. And see which list has the most occurrences.

show 1 reply
actionfromafartoday at 1:36 PM

Sounds like the best approach. :)