logoalt Hacker News

bigwheels • today at 5:43 PM • 1 reply • view on HN

> The monitorability of frontier models is degrading.

Is there more information about why this is happening? Is political pretext because it's what the labs actually secretly want, or is there a real underlying reason this is unavoidable?


Replies

cleverpotato479 • today at 5:49 PM

Chain of thought tokens are vectors that have the same dimension as the input/output embeddings. This allows them to be un-embedded back into text, making interpretability easier.

There is no mathematical reason that the chain of thought couldn't happen in a different dimension. Indeed there are likely many reasons to do so. At this point you'd have to do some kind of (potentially lossy) projection back into the embedding dimension in order to understand what's happening.