logoalt Hacker News

clhodappyesterday at 5:11 PM3 repliesview on HN

Seems like they are closer to scratch than reasoning... Generating some scratch to draw from helps make it easier to compute the real answer.


Replies

forgotTheLasttoday at 1:38 PM

That's my personal theory too. The model is stuffing its own context with vaguely related tokens, which helps the attention heads retrieve the right tokens.

show 1 reply
throw310822today at 8:22 PM

It's also interesting because in humans the existence of "Aha!" moments that are not preceded by or are only loosely related to a chain of thought is taken as the proof of the fundamental mystery and irreproducibility of human intelligence. Now the same argument is made to deny that LLMs actually think. Go figure.

show 1 reply
cyanydeezyesterday at 7:21 PM

I assume theyre searching the local gradient to see if theres a better descent before proceeding.

show 2 replies