logoalt Hacker News

Gecko4072today at 6:52 AM2 repliesview on HN

Can’t this be extended quite far? Use a cerebras-served model, use verification techniques to generate and solve millions of problems and then use that as training?


Replies

kevincoxtoday at 11:52 AM

This isn't latency bound, it is trivially parallelize. So you want to run it on the most efficient compute you have, not the fastest.

show 1 reply
gvkhnatoday at 8:15 AM

That’s the whole point, just cost and compute limitations in your way (mostly).