Was this an out of sample test of the router or was it trained on these specific use cases/eval suites?
Neither. They ran all the cases using both models, picked the winning result for each one, and then said "if you had a router that guessed with 100% accuracy, here's what it would have picked"
Neither. They ran all the cases using both models, picked the winning result for each one, and then said "if you had a router that guessed with 100% accuracy, here's what it would have picked"