a hile ago (when big providers still provided logprobs) i created a VS Code highlighter that visualizes unsure tokens.
Since most chat models want to answer with a human-readable message i think their logprobs are not as meaningful. It would be interesting to see if one choice is like "correct" and if the model wants to choose it more often, cause it might not answer the question but to prose to the user.