logoalt Hacker News

qayxc • yesterday at 12:24 PM • 1 reply • view on HN

When put to the test in real-world environment, the capabilities don't look as impressive as benchmarks and synthetic tests might indicate. So doubts about actual spatial reasoning capabilities remain.


Replies

simianwords • yesterday at 12:50 PM

I see. Gary Marcus said that AI won’t be able to make a coffee in any arbitrary home kitchen.

I think it’s a good test and I think LLMs will reach it in 3 years. Current benchmarks maybe slightly incorrect.

I’m happy to make a 4:1 bet in my favour that I’m correct about the kitchen bet.

➕ show 1 reply