It might be beneficial while not being optimal on its own.
The obvious example is if it has different behaviour around local minima, it could be an altenate pathway out.
I have often wondered if doing training with radically different aproaches for the first few iterarions would avoid any method specific artifacts before the weights had time to denoise.
You can probably distribute training more easily too
Yes, it's interesting research, I agree and even wrote so in my post :-)
But the submission title calls it a "backprop alternative." It is not, at least not yet.