logoalt Hacker News

prometheus1992yesterday at 10:34 PM2 repliesview on HN

Does this mean they ended up sharing those private codebases with OAI, Anthropic etc? Also, the ~30% number tracks with my experience. I thought I was going insane for expecting too much from the models but they are still bad, including astra. This morning it messed something pretty trivial while fixing an issue which I was shocked to see. Also2, benchmarks don't mean much these days.


Replies

ramigbyesterday at 11:28 PM

Can you please share, if you are comfortable of course, what did the model(s) mess up? what were you using codex/cc/pi? did the project have a solid agent.md/claude.md? I am genuinely curious whenever someone have such a low success rate with models what is happening because it could be fixed maybe? From my own experience using agents for the past year or so. The rate if I have to guess, is well above 70%. I mainly use claude (opus) on typescript react projects that are well setup with minimal plugins/MCPs!

happy to share more if you are interested.

show 4 replies
doctorpanglosstoday at 3:32 AM

the requests went into a pipeline that turns them into de-identified, but salient, training data, yeah. everywhere except maybe bedrock.