A lot of people are posting here about how bad the end product is, but that is kind of the point. Models have moved beyond generating images to a new kind of benchmark that better exposes understanding of the physical world, and we can use benchmarks like this to measure future progress. (Of course, it will have to be a qualitative/subjective measurement.)
I worked with an LLM to build a ~3D animation of the Back to the Future delorean Time Machine as a way to spice up the hero on a docs page.
That took a fair amount of custom tuning and I had to create a tuning view to get some of the behaviors right.
But it was enough fun that I generalized it to take in ~any scene description from a film. It goes out and gets more detailed descriptions and film stills if available but also takes custom stills if you provide them.
My test scene was the Gauntlet scene from Apocalypto. It is low fidelity but does a pretty amazing sequence with somewhat believable physics of the javelins etc.
Here is the docs page with the vertical takeoff / 88 miles an hour time travel: https://contextify.sh/docs
I can share some of the Apocalypto bit if anyone is interested.
It seems pretty clear that Anthropic models have been specifically trained to be good at generating three.js (JavaScript 3-D Graphics) code, so given current state of AI code generation in general, I don't find three.js models/animations as indicative of anything other than the model's ability to write three.js code.
When Fable was first released the day-1 demos of it on Twitter (presumably from people who were given early access, and/or Anthropic employees) were pretty much 100% three.js stuff. Yes, it looks nice, but it doesn't tell me any better than an Erdos proof whether the LLM will be able to run my vending machine.
I'd rather have them battle on the topic "Who builds a better Google Wave for LLM chats" to explore the space of how AI studios could be.
"Getting Started with Google Wave": https://www.youtube.com/watch?v=eKUAqNGVwX0
this is not a good benchmark for models, but it's great if you're optimizing for attention on twitter because video content and 3d animations perform best on social media.
a real benchmark is instead running evals on your own traces, and building a cost/quality/speed profile for models based on real workloads. but it doesn't get you a shiny video you can post on twitter.
I can forgive the modeling being godawful jank (windows floating in the air, disconnected from the house). But I expected it to have a better understanding of the text. Instead, we have Bilbo's "disappearance" interpreted as him magically transporting or cloaking, and similarly for his reappearance.
We just need to go one layer deeper:
Generate an SVG of an ai generating an SVG of a pelican on a bicycle.
https://chatgpt.com/s/t_6a6fc59ae00081918095322a63e1503cI always thought of the Pelican more of like a gimmicky quick test. There are people who took it as a serious benchmark for overall model performance?
I think the most important insight is the limitation of LLM perception:Slowly taking screenshots.
That method of perception probably scales N^2... so sure with more compute, LoTR animation will improve. But I think to get a real jump in "experiential feedback", perception needs to scale linear or sublinear. Maybe that's there LeCunn's jepa will come in.
There needs to be the removal of the middle man:
image -> text -> action
To image -> action.
Reading the title, I was thinking "Karpathy? I don't know this chess player." (The Pelikan is a well known chess opening, and famous chess players often have book titles like "X's Y" where X is the player, and Y is the opening)
I’d like to see a human one shot a pelican on a bicycle in raw svg.
Speaking of benchmarks has anyone given AIs Where’s Waldo pages and asked it to find Waldo?
I’ve been trying it on them all and can’t find one that does it consistently. The best will tell me they can’t. The worst confidently point out one of countless Waldo-likes.
How long we will be testing and benchmarking generative capabilities of LLMs? in creativity, in code generation less or more result is expected and approved, but in execution $19.8 can not be $19.9 or $19.7.
IMO, the area where AI is going to be most useful over the next couple years is in developing manufacturing processes top to bottom. Maybe a million token budget is too small, but something like "design me a sneaker and all the equipment to manufacture it autonomously".
It would be interesting to see the models work on a book it hasn't been trained on yet. I guess sadly that means any book released very recently.
Definitely impressive demo, I do wonder though if the countless artwork, films, images etc produced over many decades around Lord of the Rings somewhat influenced the outcome of this though.
How much are the hobbit houses described in the book? The ones here look exactly like the movie
Regarding the argument about LLMs having difficulties auditing their work:
I wonder whether we are entering the era of throwaway software. Just like cheap plastics and improved processes has enabled us to rapidly manufacture anything we want for a very low price, maybe LLMs give us the same for software. Produce it cheaply and if it breaks throws it away and reproduce it.
I'm pretty tired of the "Y made this game in Z tokens" all over the internet last week. They look impressive, and it's cool that it's even possible, but they're useless as games. None of them are any fun. They're like the most boring variant of basic controllers you can imagine. None have any cool mechanics. None have any tweaks made from hours and hours of testing. All have the same cel-shader.
There is a tipping point between procedurally generating everything in SVG to maybe giving them tool access to something like 3dsmax (or having them build and then use a tool to do the thing vs doing the thing).
We must not be using the same opus 5, because if I tried to generate this it would refuse based on copyright grounds.
Is anyone else getting "mongodb is webscale" vibes? (Except 16 years ago that was a lot smoother, because it used some sort of "render this conversation" engine...)
This has started to feel a bit like the beginning of railroads and then the steampunk fiction of "let's just build railroads to everywhere". We don't need it and there's no use for it.
As with painting, after a while there's nothing really new to paint, we genuinely need 0 new software. We need to fix our broken physical world, our social lives, our kids and what's left of our democracies.
This software crap is done, leave it to the nerds.
I suspect LotR is a singularly unrepresentative choice here considering how much info exists about it.
I'd like to see the Silmarillion, specfically both Ainulindalë and the Fall of Numenor. At this point a visual model would probably produce something better than Amazon (but presumably not Jackson).
I really dislike this AI programming thing of “Mr LLM, go slam your face into the problem until there’s no problem left, then call me back”. (Not sure if it’s a recent trend or a fundamental nature.)
It always brings to my mind some words from Rich Hickey:
I think we’re in this world I’d like to call “guardrail programming”. It’s really sad: we’re like, “I can make change because I have tests!”. Who does that? Who drives their car around, banging against the guardrails, saying “whoah, I’m so glad I have these guardrails so I can make it to the show on time!”
I don’t think I really have a point to make here, other than it just feels like someone’s released a bunch of carnival bumper cars onto the highways.In my experience SVGs are still too hard for LLMs.
I gave Fable a jpeg and asked to draw an SVG, using a loop that renders the SVG into an image so Fable can inspect it.
Results looked like drawing of a 5 year old.
One difference for human to understand the video is that we only care the changes on a picture compare to LLM
On consumer devices, AI communicates with us through speakers and screens. Screens are the richer medium, so most consumer AI innovation will happen there.
Computer graphics will have enormous applications because they are directly controllable by LLM-generated code. Video models are probabilistic and less suitable when precision matters. In education, for example, we need exact visuals. If an AI wants to plot y = sin(x), it should generate the precise graph through computer graphics rather than approximate it with a video model.
Still images seem like a better quick test because we can see them at a glance. Maybe ask it to make a comic?
I think I could tolerate 50 Shades of grey rendered in this style.
This makes me think about using a game engine and a coding agent instead of current video generation AIs. It will probably cost much more, but it will have almost zero consistency problems. Is this line explored?
The tilde thing remains uniquely obnoxious in a field that seems want to mangle language for fun, so that's innovative I guess
> I also like this kind of examples because no one in their right mind would ever spend the time to write something this custom
This is an odd take, given that Karpathy is certainly aware that the LotR films absolutely did create Bag End in digital format; that their creation was outstandingly high quality; and that Claude’s output here very obviously “leans heavily” on their prior art.
Do you guys notice that LLM can create fancy viz/animations by coding them instead of leveraging what we humans usually use (e.g Lottie, After Effects)?
I wonder if Flash is still popular... LLM can use that instead...?
Anthropic spokesman [0] Andrej Karpathy is here to tell you about token-wasting loops, and insists on the weird idea that they are "~free", when in fact, they are fuelled by expensively burning investor money.
[0] Seriously. Get used to mentally prefixing his and Boris Cherny's name like this, every time you see them quoted. These people are speaking while employed; there is no chance they are not aligned with the employers who will make them wealthy. The tech industry does like to pretend that for some reason AI people, uniquely, speak thoughts unbiased and for themselves or even for science or humanity.
Pelican maxxing.
I love how nobody cares about copyright anymore. Not even an after thought. Might be okay if you're an employee of anthropic/openai but for us mere mortals I'm not sure I would share something this blatant. In US it's $150k+ per violation and you've given them all the proof (even a confession)
you don't really need screenshots if you have an engine expressive enough for the scene generation while ensuring the visual appearance of the engine output itself is feasible.
that's why these things are actually pretty good at openscad/freecad/F360 mcps , the visual reality is enforced and guaranteed by rigor in the interpretation engine that is anchored to human physical reality.
I'm sad that Andrej Karpathy went from being one of the most reasonable, trusted, and credible voices in AI to peddling marketing slop for Anthropic.
8 months ago, he was (very reasonably) claiming that reliable agents are at least a decade away, but this now goes against the interest of his employer, so the narrative has been changed.
I can't believe the video demo is $10
> I also like this kind of examples because no one in their right mind would ever spend the time to write something this custom
There are people in their right mind who would do that and their are already examples of people who did similar things.
But maybe not in the future if people would confuse all the effort with AI
it shows the possibility that SVG replace PNG/JPG even video.
After watching this video I am absolutely positive that the issue is not a lack of stamina in humans. It is that humans have the capacity to realize that this is a bad idea long before they complete it.
It may make sense to switch this to USD/Omniverse.
> I gave it the first paragraph of the Lord of the Rings, a 1M token budget (~$10) and asked for three js render of it.
I think it's interesting that the "Bag's End" interpretation in the video clearly looks like the one from the movies, but generated here as a three.js 3D asset.
It makes sense that the movies (or shots/frames from them) were in the training data, and I can also easily imagine an association in concept space between the textual description of Bag's End and the frames from the movie.
But how on earth does the model then go on and convert the latent representation of those images into coordinates for a 3D mesh, without ever even restoring the image? In what kind of representation are the images from the movies stored that it can do that?
AI is now somewhere between "Money for Nothing" (https://www.youtube.com/watch?v=wTP2RUD_cL0) and "Knick Knack" (https://www.youtube.com/watch?v=9uhM_SUhdaw) in capabilities.
If you like this check: https://news.ycombinator.com/item?id=47400868 it can be used to generate animations (not games) as well.
> sure, why not, it's ~free
Yeah, please check in with the folks protesting data center builds in their town causing their electricity prices to skyrocket and tap water to turn into a scarce resource.
And now, instead of actually doing the above, please go ahead and downvote me, because how dare he question those LLM games.
I don't think it's a bad way to benchmark new models, I just find it concerning that the author implies that "pelican on a bicycle" has been exhausted. At the risk of making overly broad, unfalsifiable claims I think multi-year exposure to AI content has dramatically raised our expectations for speed and volume but lowered them for quality. We see a very janky pelican and declare the problem solved.