logoalt Hacker News

prohoboyesterday at 8:40 PM1 replyview on HN

Both of your comments are illuminating :p

So, we could technically debug a prompt's output? I get that there are too many steps to actually step thru, but what if there were checkpoints? At least you could isolate behaviors to specific sections of a neural network?


Replies

bonoboTPyesterday at 9:06 PM

Of course. And mechanistic interpretability research is a thing.