logoalt Hacker News

joemitoday at 3:29 AM1 replyview on HN

I think you're both saying the same thing: the digitization of the article wouldn't have been the source, since 彁 would have had to exist before the digitization happened in order for the OCR to misread 彊 as 彁.


Replies

usrnmtoday at 8:52 AM

Most kanji are a combination of several smaller parts called "radicals" in English. If you look at these two kanji through this lens, you will see that it's actually a very simple mistake, one existing radical is replaced by another existing radical. It is very easy to imagine software that was working exactly like that: interpreting kanji as a combination of radicals rather than individual unrelated symbols

show 1 reply