OCR the PDF.
It's the ground truth, and it's not like it's more complex than parsing a pdf, at this point, the technology for OCR might even be better than PDF parsing, which is full of accidental instead of natural complexity.
I don’t really understand how OCRing a PDF could be more accurate than reading the text nodes.
Another thing to consider is OCR works well for English but not so well for other languages.
I don’t really understand how OCRing a PDF could be more accurate than reading the text nodes.
Another thing to consider is OCR works well for English but not so well for other languages.