Just speculating but I "feel" 4.7 was post-trained using more synthetic techniques. The wa...

b--l • yesterday at 11:17 PM • 1 reply • view on HN

Just speculating but I "feel" 4.7 was post-trained using more synthetic techniques. The way it writes for one thing, it's "personality", is less human and more fatiguing-AI-slop like.

Replies

ACCount37 • yesterday at 11:27 PM

You don't need to fry with RLAF to get that "slop feel". The first iterations of "AI slop" were raw SFT+RLHF - all human input, all inhuman output.

That said, I completely agree that 4.7 was a pronounced "model personality" regression. Closer to ChatGPT, and I mean that as an insult. Yet to check whether 4.8 is better.

alt Hacker News

Replies