The jump from 3.5 to 4 felt gigantic to me back then.
GPT 5.0 did feel underwhelming though.
Oops, I might have been misremembering then. Maybe I meant 4 to 5
GPT 4 to 5.5 felt about the same as 3.5 to 4 to me.
Yeah, GPT4 was one-shotting utilities that GPT3 Davinci couldn't. So, I'd have my limited tokens on GPT4 crank out the initial program before iterating with my abundant, GPT3 tokens.
Agree but it's helpful to remember how we were personally benchmarking. I remember people saying stuff like "haha I asked gpt4 for xyz function and the typescript didn't even compile". We're so far beyond that now, we just adapt quickly.