logoalt Hacker News

Thinking fast and slow in AI: The role of metacognition (2021)

169 points • by teleforce • today at 3:23 AM • 73 comments • view on HN

Comments

gchamonlive • today at 10:41 AM

There was this post a few days ago https://news.ycombinator.com/item?id=49797323

It had this to say in the linked post:

  This led to the natural question: can gzip do language modeling? (...). Here’s some real, unedited output after priming it on tiny Shakespeare:

  gzipt --corpus data/tinyshakespeare.txt --prompt $'MENENIUS:\n' --length 200

  MENENIUS:
  'Though all at once canq

  MARCIUS:
  Pray now, nocamest thou to a morsel.

  LARTIUS:
  Hence, and
  I' the end admire, where G
  again; and after it ag .
Now thinking back, what's missing so that gzip could unwind the correct body of work from Shakespeare is just a correct sequence of bytes. One way to arrive at this is by just getting the body of work and doing the inverse, compressing it to get that golden sequence of bytes.

The other is what thinking does, it tries to predict the missing sequence of tokens from a high entropy source, the prompt, in order to increase the likelihood of correctly decompressing the desired results from its weights.

➕ show 2 replies
red75prime • today at 5:41 AM

It reminds me of "At the time we drew boxes labeled 'perception', 'cognition' with arrows between them." An imprecise quote that I can't place.

I guess my box labelled 'subconsciousness' is trying to say that low-level mechanisms that give rise to the observed cognitive phenomena might have nothing to do with neat boxes.

➕ show 2 replies
hoppp • today at 2:27 PM

Most humans have weak meta-cognition, a large percentage doesn't have verbal thoughts.

Meta-cognition makes sense in a dynamic and updatable and modular system, for example I can monitor thoughts coming from my amygdala with my prefrontal cortex and then adjust how I process these thoughts.

In LLMs it makes zero sense, even if you feed the output of one model into another, there is no way they can update the heuristics behind how those were computed.

arbirk • today at 11:28 AM

Interestingly that is not what we got, but maybe we should loop at architectures like this again? The JEPA loop is interesting, but might fail for the in-flexibility of the component ordering

creativeSlumber • today at 7:30 AM

How relevant is this fast/slow thinking thing with regards to current frontier models?

I know a large organization who's built their AI framework completely around this concept, and I feel that it's not really meaningful concept with the capabilities of current models.

➕ show 6 replies
crorella • today at 5:31 AM

It looks like a lot like how data bases query optimizers work, with the exception that in the paper there is also a learning/memory component that conditions the evaluation of the answer provided by the first model.

➕ show 2 replies
zfoong • today at 7:13 AM

At least this is written before ChatGPT.

vist_orn • today at 8:03 AM

Trying to get LLMs to 'think about their thinking' is my daily struggle. This paper nails why it's so critical.

readthenotes1 • today at 9:04 AM

If I recall correctly, all that fast and slow business has been debunked as yet more non-replicable pop psychology.

I shouldn't be surprised that it shows up in a screed on AI

bbor • today at 4:43 AM

This is still a great paper, but it's missing the second axis of the quadric -- if the only two options are thinking fast or thinking about thinking, that leaves no room for thinking slow yet deliberately, AKA selfconsciousness. See https://www.gutenberg.org/cache/epub/4280/pg4280-images.html for details

I do wonder if any of these folks ever got a chance to try this at one of the big labs, tho...

➕ show 1 reply
jannyfer • today at 4:25 AM

> submitted Oct 5 2021

(In case people miss that before discussion)

➕ show 1 reply
locitra • today at 3:44 PM

[flagged]

ot4t • today at 2:10 PM

[dead]

aidiscoverywire • today at 6:00 AM

[flagged]

BigDogAU2026 • today at 10:42 AM

[flagged]

tug2024 • today at 3:16 PM

[dead]

simianwords • today at 6:20 AM

This has already been solved by GPT 5 Adaptive reasoning. A single model that knows when to reason or not based on a thinking parameter we provide (like xhigh). What’s the relevancy to post it today?

edit: why is this downvoted?

➕ show 3 replies
sinuhe69 • today at 5:51 AM

2021. Please remember the rule of HN to add the year if it’s not actual.