They're not magical go faster juice. I don't know of a microarch where they're faster than modern GPUs at ML training or inference.
Having the model weights in static mask ROM would massively improve power efficiency for inference (see what Taalas is doing).
Having the model weights in static mask ROM would massively improve power efficiency for inference (see what Taalas is doing).