Is the benchmark measuring one-shot retrieval accuracy, or Coding agent response accuracy?

esafranchik • yesterday at 5:25 PM • 1 reply • view on HN

Hey! Co-author here. The benchmark currently only measures retrieval accuracy.

We’re interested in measuring it end to end and also optimizing, e.g. the prompt and tools, for this, but we just haven’t gotten around to it.

➕ show 1 reply

alt Hacker News