I hope the LLM wave will leave GPUs behind to go back to pursue more general-purpose computation rather than spending their die area on multiplying 4-bit-number matrices and such things.
Is that not literally the exect opposite of the direction asics for LLM inference is going?
Is that not literally the exect opposite of the direction asics for LLM inference is going?