Why would a local model for a consumer device need 17k tok/s?
Apple is better off building chips with generalizable TPUs (or equivalent) so they can upgrade/patch models.
I cant shake the feeling of "640KB is enough for everybody". Imagine not one AI answering over 1 minute but a team of 100+ agents in hieararchical structure taking care of your request in seconds, checking each other.
I cant shake the feeling of "640KB is enough for everybody". Imagine not one AI answering over 1 minute but a team of 100+ agents in hieararchical structure taking care of your request in seconds, checking each other.