logoalt Hacker News

simonwtoday at 7:04 PM7 repliesview on HN

This is fantastic

I've been casually spot-checking other animals in other vehicles, because my absolute dream situation here is to catch an AI lab that's demonstrably better at pelicans on bicycles than other combinations.

Catching a lab cheating specifically on my one dumb benchmark would be really funny.

Dylan's methodology here - generating 1008 SVGs across an 8x6 combination - is significantly more robust than anything I was considering.

His conclusion:

> Nothing jumped out at me. I couldn’t find a case where the pelican-bicycle images looked noticeably better than the rest of that model’s grid.


Replies

lukevtoday at 7:38 PM

What if they’re not pelicanmaxxing, but svgmaxxxing in general?

Because otherwise using a LLM to generate complex svgs is pretty niche and what I thought made this a good benchmark when it was new - generalized programming and spatial knowledge.

Obviously image gen in svg format is not a particularly hard problem if tackled directly on its own.

show 8 replies
eobtoday at 8:25 PM

Simon I hope from this day hence, your bio always includes:

"Simon Willison, among other things, is an advocate for the inclusion of pelican geometry in LLM training datasets."

gopalvtoday at 7:45 PM

> Catching a lab cheating specifically on my one dumb benchmark would be really funny.

Similar thing happened when TPC came up with SQL benchmarks.

If you're not good at TPC, your engineering team is no good.

If you're good at TPC, then (as a customer) we will actually include you in a bake-off benchmark for our specific problem.

Winning on it is the price of admittance into the game, especially in a crowded market.

But how narrowly you benchmarket matters, you can't just hard-code that specific scenario & not fix anything adjacent while you're at it.

For example when it comes to GPUs, the "Quack3" (sic) benchmark on ATI cards comes to mind.

docheinestagestoday at 8:10 PM

I think a more fundamental test is SVG art creation in general. Perhaps a pipeline to take any image, caption it, ask the LLM for an SVG, rasterize to an image, and finally either use a deterministic visual similarity check or ask another LLM to be the judge and score how close the SVG is to the original image.

show 1 reply
gilleaintoday at 7:09 PM

Perhaps also vary the bird? Wikipedia tells me pelicans are in the order _Pelecaniformes_ so shoebills or herons might do.

cyberaxtoday at 8:15 PM

> I've been casually spot-checking other animals in other vehicles

Snakes on a plane, weasels on a diesel, spiders on a glider, baboons on a balloon, goats on a boat.

mattertoasttoday at 7:26 PM

[dead]