logoalt Hacker News

robrenaudtoday at 10:02 PM0 repliesview on HN

Regarding the temperature explanation:

> "Instead of picking the highest-probability token, we can use different selection strategies to balance safety and creativity in the generated text".

Safety is definitely the wrong word here.

Temperature 0 generated text actually has a weird "lack of surprise" character that makes it seem artificial. [1]

> "high-probability texts can be dull or repetitive. Humans use language as a means of communicating information, aiming to do so in a simultaneously efficient and error-minimizing manner; in fact, psycholinguistics research suggests humans choose each word in a string with this subconscious goal in mind."

I'd completely drop the dropout explanation. It's just not part of the modern recipe anymore, AFAICT.

As for the ambitious goal of explaining transformers with a single interactive visualization, I just have a hard time imagining a person is going to newly understand both word embeddings (word2vec blew my mind in 2014) and also gain an understanding of attention.

I am making my own visualizations for a presentation on "Full Bandwidth Transformers"[2] that I am giving tomorrow at the Deep Learning Study Group (SF) (on zoom for the non-locals)[3]. It's not meant to be stand alone/context free, but I'd love some feedback.

https://rrenaud.github.io/fullbandwidth_transformer_viz/

[1] https://arxiv.org/abs/2202.00666 [2] https://arxiv.org/abs/2608.08888 [3] https://www.meetup.com/deep-learning-sf/events/316601593/