turns out the bottleneck was never model size, it was having someone who actually understood the problem define the reward function
I agree very much. I'm getting the same take-away even more now that I'm using autoresearch as my main strategy.
And I wonder: We've got all these amazing (programming) languages to define solutions; where are the languages to clearly define the problems?
I agree very much. I'm getting the same take-away even more now that I'm using autoresearch as my main strategy.
And I wonder: We've got all these amazing (programming) languages to define solutions; where are the languages to clearly define the problems?