logoalt Hacker News

mattnewtonyesterday at 4:26 PM0 repliesview on HN

The exception being the token embeddings and lm head (which scale with the number of tokens the model knows and presumably you need a smaller number in the tokenizer for only English and python). But those are a pretty small % of the total model weights on most LLM sizes