You don't need to fry with RLAF to get that "slop feel". The first iterations of &quo...

ACCount37 • yesterday at 11:27 PM • 0 replies • view on HN

You don't need to fry with RLAF to get that "slop feel". The first iterations of "AI slop" were raw SFT+RLHF - all human input, all inhuman output.

That said, I completely agree that 4.7 was a pronounced "model personality" regression. Closer to ChatGPT, and I mean that as an insult. Yet to check whether 4.8 is better.

alt Hacker News