You need to train data for a BERT-based classifier, and then there's a risk that it will pick up specific biases from the data instead of what you want.
As far as I understand, the idea of Jev is zero-shot or few-shot classifier: it learns a lot of stuff at pre-training, but unlike a classic LLM it doesn't need to learn how to chat, so it can be much smarter at a particular size