logoalt Hacker News

Lercyesterday at 10:26 PM3 repliesview on HN

It might be beneficial while not being optimal on its own.

The obvious example is if it has different behaviour around local minima, it could be an altenate pathway out.

I have often wondered if doing training with radically different aproaches for the first few iterarions would avoid any method specific artifacts before the weights had time to denoise.


Replies

cs702today at 9:48 AM

Yes, it's interesting research, I agree and even wrote so in my post :-)

But the submission title calls it a "backprop alternative." It is not, at least not yet.

dnauticstoday at 2:32 AM

You can probably distribute training more easily too