logoalt Hacker News

BobbyTables2 • today at 2:35 AM • 5 replies • view on HN

I’m a bit behind the times…

How does Opus actually paint? Thought it only generated text…


Replies

nater5000 • today at 10:05 PM

Opus can see images, it just can't generate images.

So it can "draw" an image programmatically (e.g., "place a black pixel at (0,0), place a blue pixel at (0,1)," etc.), then look at the image it is drawing as it is drawing it.

Of course, you should be aware of the "pelican riding a bicycle" SVG benchmark test, which is effectively this process but in one shot.

vunderba • today at 5:39 PM

Think of it like the old turtle graphics feature in the LOGO programming language.

You give an LLM a virtual canvas and a set of “instructions” to control the pen.

https://en.wikipedia.org/wiki/Turtle_graphics

ncr100 • today at 5:51 PM

Another idea to make Claude generate images:

Ask generative AI to create an embedded web page of a voxel Mona Lisa head. It will generate an image.

Ask it to sample that image and extract out of it as binary that is well formed in the PNG data format...

EDIT: https://claude.ai/artifact/HUUtPMp7aF8iyjimViBSor - generates the Mona Lisa as a PNG via Claude sonnet 5.5 medium level. It's not a very good Mona in my opinion. Fire it up, rotate the 3d voxel presentation, scroll down, click the button, scroll further down to see the PNG. Here is the chat with prompts: https://claude.ai/share/2fb3126e-e97a-4a1a-9571-51d16f054330

TeMPOraL • today at 5:53 PM

Half of the industry is, going by the amount of "AI doesn't amount to anything and is just regurgitating text".

Two things:

- For the past year or more, many models can emit images directly, and all models that matter can see images directly - that's what "multimodal" means. Tokens don't have much to do with textual language anymore, they're more like units of sensory experience.

- Even restricted to text, a language model can operate anything that can be expressed as text, as long as you have a translation layer between textual representation and the final form. That includes giving commands as text. The total addressable space of what models can be used for is, thus, approximately anything humans do.

➕ show 2 replies
llagerlof • today at 4:22 AM

It generate commands and coordinates to move the mouse and use the paint tools.