logoalt Hacker News

disgruntledphd2today at 7:57 AM2 repliesview on HN

Wasn't this during the period where they had a bunch of bugs around caching and the models were making loads of weird decisions?

I honestly feel like basically nobody knows anything about these models, it's all just vibes (and I'm no different).


Replies

Bluesteintoday at 9:12 AM

Indeed.-

troupotoday at 8:13 AM

Anthropic broke their models in spring, denied it, gaslighted everyone who said so, and then all but admitted it: https://www.anthropic.com/engineering/april-23-postmortem (basically doing Anthropic things).

> I honestly feel like basically nobody knows anything about these models, it's all just vibes

This, too. Since only providers know what they actually serve, what they change and what limits they impose.

There are some visible degradations though. E.g. Claude-ish.

As for a personal anecdote: around February I created a rather complex quiz web app for myself and friends with multiple question types, sync between screens, multiple media upload types, multiple scoring and timing types, MC inetrface etc. etc. etc. It took me a week or so in the evenings with rather vague prompts to make it.

Now Claude (and Codex) cannot reliably build a much simpler web app even with precise instructions while also maintaining the visual consistency.

But I will agree with you, it's a feeling, not a precise measurement.

show 1 reply