logoalt Hacker News

jjcmyesterday at 6:06 PM18 repliesview on HN

Here's a image->html test. Gemini has always swung above its weight class for vision work, so I'm always eager to try it with this.

Original images: https://image.non.io/neonRamenDesigns.webp

Gemini 3.7 build: https://html.non.io/neonRamenGemini3.7

Opus 5 build for comparison: https://html.non.io/neonRamen

Opus is still best in class for this, but it's worth noting how well Gemini 3.7 does vs a more comparable LLM price wise, which is Grok 4.6: https://html.non.io/neonRamenGrok4.6 . I thought Gemini would blow Grok out of the water (it generally has in the past), but Grok has really caught up.


Replies

jjcmyesterday at 6:17 PM

Other thoughts: I really think Google has fallen behind here. Even as a high speed offering (this build took ~7min, which is pretty good!), it wont be able to claim dominance for long with cerebras announcing the Sol preview today: https://www.cerebras.ai/blog/accelerating-gpt-5-6-sol-ultraf... .

It's not a bad model by any means, but I just don't know what situation I'd reach for 3.7 Flash first for. Google really needs a differentiator, especially given how hard it is to get an API key from them. They can't be high friction and non-pareto.

show 8 replies
butliketoday at 5:00 PM

There's one specific aspect I like better with the grok version: The prices are above the fold. ON the Opus and Gemini versions, I have to scroll to see the full menu item showcase.

a2ff6eeb0today at 1:52 AM

These look exactly like all of the low budget bodega signs near me. They also look like a bunch of cheap ads for parties that I keep seeing. The sameness of style is uncanny.

(I don't have the Bodega signs, but I'm thinking of shit like this, from a quick google: https://linkstub.com/en/wet-wild-foam-party)

show 2 replies
orliesaurusyesterday at 6:55 PM

I think both outputs are really good. I don't see a lot of differences. So what exactly should be looking at and notice that one model did worse or better than the other one.

EDIT: OKAY I see it's mostly the "image" generation, not so much the HTML... Noticeable in the food photos and the foodtruck/cart photo

show 1 reply
flockonusyesterday at 6:51 PM

There is some irony being a developer and reading along the lines of: "oh look at the comparison between these models executing a task for a few cents on a job i'd be charging 1k minimum"

show 3 replies
codazodayesterday at 6:21 PM

How are you doing this with Opus. Clearly I’m missing something. I always turn to ChatGPT when I need images because Opus typically refuses. I’ve tried Claude Code and Claude online in the past. I’m pretty sure neither created images for me and I thought this was because Anthropic was focused on code.

I guess I need to try harder. :)

show 4 replies
keyleyesterday at 11:37 PM

FYI Gemini's version is less broken than Opus' in Safari...

jzemeocalatoday at 2:34 AM

One of my favorite image tests with AI models is schematic analysis...I build and repair tube amps for a living, and use AI for such work a LOT.

so far, IMHO, the best has been opus and fable\mythos.

show 1 reply
joshmnyesterday at 11:27 PM

What’s the prompt you used for this?

Edit: oh wow, diffui looks nice!

snissnyesterday at 6:13 PM

I'm curious how much the harness plays into this. I'm somewhat surprised by the gemini and grok results, they seem to have strongly deviated from the original images. I'm thinking maybe the harness has a big effect? It's possible to proxy in different models to claude code, if you're curious you might find it interesting to test!

show 1 reply
cechmastertoday at 6:48 AM

I'm not sure what you consider good design, but if it's subjective, then I see it differently from your examples.

Gemini 3.7 looks the best. Opus 5 looks almost as good as Gemini. Grok 4.6 looks pretty terrible.

getnormalitytoday at 7:03 AM

IMO Gemini's is better than all the others, including the original.

mediumdeviationyesterday at 6:37 PM

I'm not sure what prompt you put in but did Gemini replace the all of the images in the original with its own? That would be really weird behavior unprompted.

show 1 reply
giarcyesterday at 9:36 PM

Can you share what your prompt was for that?

show 1 reply
XCSmeyesterday at 6:56 PM

They both have horizontal scroll on mobile...

show 1 reply
pphyschyesterday at 9:58 PM

Was the original concept generated by Claude somehow? It gives me Claude UI vibes with all the extraneous small-caps text elements.

t3hTaoyesterday at 9:03 PM

[dead]

xyzsparetimexyzyesterday at 8:36 PM

This has so much less character than the pelican smdh..

Plus is the ramen in HK even any good?

show 1 reply