logoalt Hacker News

postalcoder • today at 1:57 PM • 10 replies • view on HN

Since this is for the mac you really should be using apple's vision framework for OCR. It smokes tesseract in both speed and accuracy.

Edit: I'm curious which LLM was used to generate the code. I fed the title of your post to claude/deepseek/qwen/codex asking to recommend a stack for this project, expecting to frown thinking that they still recommend tesseract. However, I found that they all recommend apple's vision framework. In fact the latest model to recommend Tesseract is gpt-4.1.


Replies

vavkamil • today at 3:40 PM

I recently tried to recover text from 8 frames of an office-shot YouTube video, where only a small, blurry portion of a computer screen was visible. After spending half a day with Astra on it, the conclusion was that it’s not possible to read.

Later that evening, I just paused the YouTube video on my phone, circled the part of the display with Google Lens, and it read the whole thing with pretty good accuracy. It was mindblowing :)

➕ show 1 reply
darepublic • today at 5:57 PM

My own experiments with tesseract were very mixed. I did use just a locally runnable version from a public repository. Text from webpages. Sometimes it would do great but small variations could make it fail completely. Paddleocr on single line text for me has accuracy in the range of 95%

robotmay • today at 3:25 PM

Mostly unrelated but fun thing I discovered earlier this year with Apple's vision - if you have text both correctly oriented and upside down in the same image, it likes to interpret the upside down text as a Cyrillic alphabet. I was trying to use it to read the text on camera lenses and it came up with all sorts of bizarre interpretations. If anyone's interested, I got around it by splitting the text at a point and unrolling it into a straight line before running OCR on it.

amelius • today at 10:30 PM

I mean, anything built in the AI age will smoke Tesseract.

It is a tool from a bygone era.

schainks • today at 7:53 PM

This. Learn all the knobs and dials, too. Useful performance gains to be made by tuning the right settings.

spiderfarmer • today at 2:44 PM

So the prompt was probably something like: build me x using Tesseract.

daveguy • today at 2:43 PM

nullsanity got downvoted into oblivion, but they are correct. This is one of the many reasons why vibe coding produces worse software. The code that is generated and the best practice recommendations are completely separate. They both come from a distribution of "most common", and best practice is rarely common. Especially when a practice is first established, or in a specific niche.

➕ show 6 replies
allenleee • today at 3:42 PM

Hey author here! :) Great point.

You're totally right. Apple Vision is generally way faster and more accurate on Apple Silicon. Pure Mac-only, I'd use it. (also smaller)

Main reason I went Tesseract: I want SCM to stay portable easily. It's ARM Mac for now, but the inference layer is all JS end-to-end — Transformers.js + ONNX for CLIP/SigLIP + Whisper, Tesseract.js WASM for OCR, all in plain Node workers.

A bit of background: I'm actually an iOS/macOS dev and I really love SwiftUI and AppKit — I just wanted v1 to stay portable by construction. Exploring a native Swift + MLX v2 track separately for speed.

➕ show 1 reply
nullsanity • today at 2:30 PM

And this is why vibe coders always make inferior software.

➕ show 1 reply
WokeUp420 • today at 7:24 PM

Turns out it's possible to build something without a glorified next-character search engine afterall