Here's a image->html test. Gemini has always swung above its weight class for vision work, so I'm always eager to try it with this.
Original images: https://image.non.io/neonRamenDesigns.webp
Gemini 3.7 build: https://html.non.io/neonRamenGemini3.7
Opus 5 build for comparison: https://html.non.io/neonRamen
Opus is still best in class for this, but it's worth noting how well Gemini 3.7 does vs a more comparable LLM price wise, which is Grok 4.6: https://html.non.io/neonRamenGrok4.6 . I thought Gemini would blow Grok out of the water (it generally has in the past), but Grok has really caught up.
There's one specific aspect I like better with the grok version: The prices are above the fold. ON the Opus and Gemini versions, I have to scroll to see the full menu item showcase.
These look exactly like all of the low budget bodega signs near me. They also look like a bunch of cheap ads for parties that I keep seeing. The sameness of style is uncanny.
(I don't have the Bodega signs, but I'm thinking of shit like this, from a quick google: https://linkstub.com/en/wet-wild-foam-party)
I think both outputs are really good. I don't see a lot of differences. So what exactly should be looking at and notice that one model did worse or better than the other one.
EDIT: OKAY I see it's mostly the "image" generation, not so much the HTML... Noticeable in the food photos and the foodtruck/cart photo
There is some irony being a developer and reading along the lines of: "oh look at the comparison between these models executing a task for a few cents on a job i'd be charging 1k minimum"
How are you doing this with Opus. Clearly I’m missing something. I always turn to ChatGPT when I need images because Opus typically refuses. I’ve tried Claude Code and Claude online in the past. I’m pretty sure neither created images for me and I thought this was because Anthropic was focused on code.
I guess I need to try harder. :)
FYI Gemini's version is less broken than Opus' in Safari...
One of my favorite image tests with AI models is schematic analysis...I build and repair tube amps for a living, and use AI for such work a LOT.
so far, IMHO, the best has been opus and fable\mythos.
What’s the prompt you used for this?
Edit: oh wow, diffui looks nice!
I'm curious how much the harness plays into this. I'm somewhat surprised by the gemini and grok results, they seem to have strongly deviated from the original images. I'm thinking maybe the harness has a big effect? It's possible to proxy in different models to claude code, if you're curious you might find it interesting to test!
I'm not sure what you consider good design, but if it's subjective, then I see it differently from your examples.
Gemini 3.7 looks the best. Opus 5 looks almost as good as Gemini. Grok 4.6 looks pretty terrible.
IMO Gemini's is better than all the others, including the original.
I'm not sure what prompt you put in but did Gemini replace the all of the images in the original with its own? That would be really weird behavior unprompted.
Was the original concept generated by Claude somehow? It gives me Claude UI vibes with all the extraneous small-caps text elements.
[dead]
This has so much less character than the pelican smdh..
Plus is the ramen in HK even any good?
Other thoughts: I really think Google has fallen behind here. Even as a high speed offering (this build took ~7min, which is pretty good!), it wont be able to claim dominance for long with cerebras announcing the Sol preview today: https://www.cerebras.ai/blog/accelerating-gpt-5-6-sol-ultraf... .
It's not a bad model by any means, but I just don't know what situation I'd reach for 3.7 Flash first for. Google really needs a differentiator, especially given how hard it is to get an API key from them. They can't be high friction and non-pareto.