And what do you think Jev confidence score is?
Here's a hint: confidence is not generated by a model.
Maybe I'm missing something, but why couldn't it be generated by the model? In older classification tasks with transformers like BERT, you could absolutely obtain a confidence score.
Thanks, fixed my understanding!
Do you think though that Luna being a model post-trained for chat produces over-confidence in logprobs?
Maybe I'm missing something, but why couldn't it be generated by the model? In older classification tasks with transformers like BERT, you could absolutely obtain a confidence score.