When put to the test in real-world environment, the capabilities don't look as impressive as benchmarks and synthetic tests might indicate. So doubts about actual spatial reasoning capabilities remain.
I see. Gary Marcus said that AI won’t be able to make a coffee in any arbitrary home kitchen.
I think it’s a good test and I think LLMs will reach it in 3 years. Current benchmarks maybe slightly incorrect.
I’m happy to make a 4:1 bet in my favour that I’m correct about the kitchen bet.
I see. Gary Marcus said that AI won’t be able to make a coffee in any arbitrary home kitchen.
I think it’s a good test and I think LLMs will reach it in 3 years. Current benchmarks maybe slightly incorrect.
I’m happy to make a 4:1 bet in my favour that I’m correct about the kitchen bet.