logoalt Hacker News

singularity2001yesterday at 9:50 AM0 repliesview on HN

How is that different from machine learning 101 "regression"? And why don't they just put a regression or softmax head on top of a trained transformer? (or do they?)