It's an slm (small language model) due to number of parameters and it uses the same architecture as an llm, but llms have billions of parameters
Is it? What is the cut off these days in terms of number of parameters? And where do other language models such as BERT/ROBERTA which are encoder only fit?
Is it? What is the cut off these days in terms of number of parameters? And where do other language models such as BERT/ROBERTA which are encoder only fit?