I thought the multiplications during training (backpropagation) are quite computationally expensive. The method from the paper above doesn't need any multiplications.