logoalt Hacker News

knollimartoday at 11:11 AM6 repliesview on HN

Oof 800 by 800 kills a lot of use cases


Replies

johndoughtoday at 11:33 AM

Might still be fine. The most recent crop of vLLMs proactively use whichever programs are available on the system (e.g. ImageMagick or PIL) to "zoom in" by cropping subimages if they can't quite make out the details.

show 1 reply
wongarsutoday at 11:31 AM

For most use cases you can fix that in the harness. Just give the model a tool to request a crop of specific coordinates of any image it has in its context. Call the tool "zoom" and it should be intuitive for the model

Maybe there are some use cases where you need high detail everywhere at once, but for OCR of small text and the like a zoom ability should be sufficient

show 1 reply
stronglikedantoday at 4:35 PM

I don't know about a lot. Probably more like a few. I take a lot of screenshots for various reasons, and over 800 seems like I could have done a better job framing and cropping.

shadyrtoday at 11:40 AM

It might also be due to its experimental status. Wouldn't surprise me if the GA version allows for larger input. Either that or the eventual pro version.

Chnmytoday at 1:50 PM

what are these use cases?

show 1 reply
asdfsa32today at 11:18 AM

flash vs fine details. Pick one.

show 1 reply