This is an excellent idea. I was curious about using DeepSeek OCR for exactly this purpose. But a tricky question is if we could do some sort of looping or something "energy based" and use classical search to find optimal parameters (LaTeX settings) to minimize the error (pixel difference). Me knowing I would get obsessed with the second half is what's keeping me from the first half. Maybe a vision JEPA would be good. If I had API credits to burn, I'd copy paste our two comments and see how far Fable gets.