Apple has a dedicated "neural engine" which is designed as an inference NPU. Where as Google's TPU has a dual focus, both inference and training, which is a more complex design.
That TPU training part I get, but from what I have seen the NPU is rarely used for inference by LLMs. They still use the GPU, no?
That TPU training part I get, but from what I have seen the NPU is rarely used for inference by LLMs. They still use the GPU, no?