logoalt Hacker News

Show HN: LLM Attention Visualization

156 pointsby ifzyesterday at 4:59 PM24 commentsview on HN

Comments

mncharitytoday at 1:10 AM

UX report. I wished to examine attention state step by step, but I found the animation moved along too fast for that. So I tried pausing...

On Chromium/linux, pressing pause doesn't pause, instead resetting the animation to it's pre-play state - the current attention highlighting disappears. Pressing play again, restarts at the beginning. Having a commonplace "pause pauses, and play resumes" UI, could allow more time to look over state. A youtube-like slow playback 0.25? option might similarly help. Or perhaps even better, buttons for single stepping. Tnx for your work.

MCP123yesterday at 9:33 PM

This is great, thank you. I have to teach this stuff on Friday so perfect timing. It's hard to explain the attention mechanism in a way that becomes intuitive because the weighting scheme does not help much with the intuition. Having a visualization like this helps a lot. Don't move that page please since I'll link to it!

fuddleyesterday at 6:49 PM

This is great, I've read multiple books and watched videos about the attention mechanism. Now that I understand it, this is the clearest example I've seen on how attention works.

lhk931122today at 1:28 AM

Is the attention explanation of why the model tells like this? I've seen that there are many discussions about this. (Image attention visualizations were not that good I think)

scottcodietoday at 1:21 AM

You can also mine attention from image models, it's a lot of fun and very interesting.

talhaanwartoday at 6:12 AM

thanks for making it simple and visualizable

sva_yesterday at 5:50 PM

I highly question this simplistic idea of high vector magnitude = high influence.

show 3 replies
wopakyesterday at 6:46 PM

neat, combining info from two phrases is hard to see without such a tool.

are you worried later-layer attention gets drowned out by earlier layers just because there are more of them contributing to the sum?

show 1 reply
asd000hhtoday at 3:43 AM

How it works?

itsnasmeyesterday at 6:34 PM

I like the visualisation. Pretty cool

staredyesterday at 6:49 PM

I am curious what's the actual formula.

I mean, there so many headers and layers, it is tricky to make a choice that will resonate with our intuition . Is it some weighted average? Or maybe ablation test?

show 2 replies
ex-aws-dudeyesterday at 8:05 PM

I don't know much about LLMs but does that mean you have N^2 computation with the context size since every token needs to track how it relates to every other token?

show 3 replies
fermlon30000yesterday at 10:40 PM

INSANE

mncharitytoday at 12:43 AM

[dead]

colophontioyesterday at 5:33 PM

[dead]

Yyylovyesterday at 9:16 PM

[dead]