logoalt Hacker News

RomanKornevtoday at 12:32 AM2 repliesview on HN

> key question is not just force balance: it’s whether the inextensible

This is definitely not astra, taking a guess this is 5.6, perhaps not even Sol, which does not reflect the state of the frontier (what the research was about).

And yes, the paper is already outdated


Replies

emil-lptoday at 7:15 AM

This is getting ridiculous. Is 5.6 not considered good enough to conduct experiments with anymore? Two months ago, if you weren't using 5.6, you were doing it wrong.

Now, you can't conduct experiments using the default model?

amlutotoday at 1:28 AM

> And yes, the paper is already outdated

The paper is about the fact that the benchmark’s evaluator is prone to egregious incorrect rejections of what answers that it should accept. A new model will not invalidate that issue.