logoalt Hacker News

YmiYugyyesterday at 7:43 PM7 repliesview on HN

I don't think it's a bad way to benchmark new models, I just find it concerning that the author implies that "pelican on a bicycle" has been exhausted. At the risk of making overly broad, unfalsifiable claims I think multi-year exposure to AI content has dramatically raised our expectations for speed and volume but lowered them for quality. We see a very janky pelican and declare the problem solved.


Replies

jonas21yesterday at 8:42 PM

I don't think he's claiming it's been exhausted. It's just that things have progressed to a point where people are arguing over the finer points of which pelican looks better -- which is often a matter of taste, and an indication that we've hit the knee in benchmark where models are no longer failing in obviously awful ways.

show 1 reply
BobbyJoyesterday at 8:59 PM

I think it's touching the limit of what one can reasonably expect any intelligent thing to produce with the only direction being "produce an svg of a pelican riding a bicycle".

When you aren't sure if an LLM can write an svg well, or that it will be able to form a pelican shape, or animate a bicycle, it's a good test. After that, it's all judgement: how detailed should the pelican be? pelicans are the wrong shape for a bicycle by default, so how much can I change its physiology to match using a bicycle before it isn't a pelican? Do I care about how well the client is able to render complex geometry?

It's not that there isn't room to do better, or that it doesn't tell you anything at all, but rather we've reached a point where what it tells us isn't very clear anymore.

show 2 replies
Demiurgeyesterday at 8:55 PM

Also, one of the advantages of the pelican test is that you can evaluate it all at once. There’s no reason the two-dimensional depiction can’t be made more challenging. Yes, at some point the pelicans might approach the subjectivity of a fine art painting, but we haven’t even seen a depiction that’s competent by the standards of a high school art class. That’s not to say the elementary school–level SVGs aren’t amazing - rather, I agree with your point.

dlluyesterday at 9:38 PM

Agreed. It's far from solved. Modern LLMs still generate pelican bike SVGs with obvious errors:

* some omitted the bottom of the diamond which connects from the pedals to the rear wheel

* some added an extra connection from the pedals to the front wheel, making it impossible to steer

* none could align the head tube with the fork

* none added a correct offset to the fork

* none could generate the chain properly in a way that attaches to the two sprockets correctly

I mean just look at these:

* Grok 4.5: https://s3.eu-west-1.amazonaws.com/images.dylancastillo.co/p...

* GPT 5.6 Terra: https://s3.eu-west-1.amazonaws.com/images.dylancastillo.co/p...

* Sonnet 5: https://s3.eu-west-1.amazonaws.com/images.dylancastillo.co/p...

from https://dylancastillo.co/posts/pelicanmaxxing.html

show 1 reply
meander_wateryesterday at 9:41 PM

Can someone explain what the pelican on a bicycle tests exactly? And why is it so important? I've never understood how it could translate to a useful task in real life.

show 2 replies
twostorytoweryesterday at 8:27 PM

100%. Why waste the tokens to render Lord of the Rings when the pelican test still clearly benchmarks so well.

mattmanseryesterday at 9:36 PM

Humans are drawing pelicans riding bicycles now. Just google it and you will find 5 or 10 of them in the first few results. Including a t-shirt design.

So it's a pretty much pointless test now.

show 1 reply