logoalt Hacker News

docheinestagestoday at 5:02 PM1 replyview on HN

Are the benchmarks comparing just the models while keeping the harness the same (the open-source Toast harness)?


Replies

breadislovetoday at 5:05 PM

yes for the retrieval benchmarks. For officeqa pro v2 we used Codex (as databricks did) and for Harvey LAB we used the vanilla harvey benchmark. For these benchmarks we added minimal tools to use mixedbread search and toast 1.