I wonder when/if we will start to see Apple include a silicon encoded model into their chips. Similar to Taalas build Llama 3.1 silicon with 17,000 tok/s inference.
So could the M7 actually include an AFM 3B model, alongside a generic neural engine?
That model would be larger and more expensive than the entire M7 chip.
Why would a local model for a consumer device need 17k tok/s?
Apple is better off building chips with generalizable TPUs (or equivalent) so they can upgrade/patch models.