> key question is not just force balance: it’s whether the inextensible
This is definitely not astra, taking a guess this is 5.6, perhaps not even Sol, which does not reflect the state of the frontier (what the research was about).
And yes, the paper is already outdated
> And yes, the paper is already outdated
The paper is about the fact that the benchmark’s evaluator is prone to egregious incorrect rejections of what answers that it should accept. A new model will not invalidate that issue.
This is getting ridiculous. Is 5.6 not considered good enough to conduct experiments with anymore? Two months ago, if you weren't using 5.6, you were doing it wrong.
Now, you can't conduct experiments using the default model?