logoalt Hacker News

dandakatoday at 12:35 PM1 replyview on HN

but for OCR there are much better suited models, I use mlx-community/PaddleOCR-VL-8bit


Replies

deauxtoday at 12:55 PM

Sometimes you intentionally want to verbatim keep "mistakes", sometimes you don't and want them to be "fixed". OCR-only models tend to only do one of those two, in VLM cases often the latter. With multi-modal LLMs you can just tell them (adherence of course needing evals/differs per model).