logoalt Hacker News

aitchnyuyesterday at 6:20 PM1 replyview on HN

Load-bearing (whoops) quirks were noticeable months back, but haven't most flagship models become predictable and reliable?


Replies

tacooooooooyesterday at 8:49 PM

it does seem to be moving in that direction. There were really specific things (large, complex json outputs) that gemini-2.5 flash was basically the only model that seemed capable of reliably for a long period. gpt-5+ has covered the usecase for us now pretty well but still evals slightly below what 2.5 could do