logoalt Hacker News

gucci-on-fleekyesterday at 9:49 PM2 repliesview on HN

> Not that hard to imagine, OCR existed back then?

How do you train OCR on a character that doesn't exist though? Even these days, OCR often confuses the 52 Latin characters, so I wouldn't expect as good of results with older technology and some 20k CJK characters.


Replies

evikstoday at 8:41 AM

It does exist? It's part of Unicode!

show 1 reply
Izkatayesterday at 11:18 PM

(Talking about Japanese here)

I don't know if this is how OCR would work here, but they're made of components called radicals. It's kind of like how letters make words, but in a two dimensional layout, and one letter was wrong and a nonexistent word was made.

For example 彁 is made of parts of 引 and 歌 (the first that popped into my mind).

IIRC there's only like 200-something radicals, though some can be hard to spot due to how they overlap or nest inside each other.