It reminds me of "At the time we drew boxes labeled 'perception', 'cognition' with arrows between them." An imprecise quote that I can't place.
I guess my box labelled 'subconsciousness' is trying to say that low-level mechanisms that give rise to the observed cognitive phenomena might have nothing to do with neat boxes.
Most humans have weak meta-cognition, a large percentage doesn't have verbal thoughts.
Meta-cognition makes sense in a dynamic and updatable and modular system, for example I can monitor thoughts coming from my amygdala with my prefrontal cortex and then adjust how I process these thoughts.
In LLMs it makes zero sense, even if you feed the output of one model into another, there is no way they can update the heuristics behind how those were computed.
Interestingly that is not what we got, but maybe we should loop at architectures like this again? The JEPA loop is interesting, but might fail for the in-flexibility of the component ordering
How relevant is this fast/slow thinking thing with regards to current frontier models?
I know a large organization who's built their AI framework completely around this concept, and I feel that it's not really meaningful concept with the capabilities of current models.
It looks like a lot like how data bases query optimizers work, with the exception that in the paper there is also a learning/memory component that conditions the evaluation of the answer provided by the first model.
At least this is written before ChatGPT.
Trying to get LLMs to 'think about their thinking' is my daily struggle. This paper nails why it's so critical.
If I recall correctly, all that fast and slow business has been debunked as yet more non-replicable pop psychology.
I shouldn't be surprised that it shows up in a screed on AI
This is still a great paper, but it's missing the second axis of the quadric -- if the only two options are thinking fast or thinking about thinking, that leaves no room for thinking slow yet deliberately, AKA selfconsciousness. See https://www.gutenberg.org/cache/epub/4280/pg4280-images.html for details
I do wonder if any of these folks ever got a chance to try this at one of the big labs, tho...
> submitted Oct 5 2021
(In case people miss that before discussion)
[flagged]
[dead]
[flagged]
[flagged]
[dead]
This has already been solved by GPT 5 Adaptive reasoning. A single model that knows when to reason or not based on a thinking parameter we provide (like xhigh). What’s the relevancy to post it today?
edit: why is this downvoted?
2021. Please remember the rule of HN to add the year if it’s not actual.
There was this post a few days ago https://news.ycombinator.com/item?id=49797323
It had this to say in the linked post:
Now thinking back, what's missing so that gzip could unwind the correct body of work from Shakespeare is just a correct sequence of bytes. One way to arrive at this is by just getting the body of work and doing the inverse, compressing it to get that golden sequence of bytes.The other is what thinking does, it tries to predict the missing sequence of tokens from a high entropy source, the prompt, in order to increase the likelihood of correctly decompressing the desired results from its weights.