logoalt Hacker News

cannedbread • today at 7:30 PM • 2 replies • view on HN

IMO by asking Jev underspecified questions like this, you're essentially using it as a random number generator (similar to the dice example). On actual NLP problems (including ones with uncertainty under human review) it does appear to be well calibrated: https://leonardgrazian.com/blog/jev-calibration/


Replies

unholiness • today at 9:51 PM

In the dice example[0] and in this one, the expected output is a distribution e.g. "1: 16.7%, 2: 16.7%...", not a random generation. I agree that in implementation, the architecture is not designed to accurately calculate or incorporate any known uncertainties like these, but the task itself is completely reasonable and arguably trivial.

IMO the lesson here is: even in trivial cases, Jev's outputs are just ~reasonableness scores which do not correspond to actual probabilities. They should not be treated as actual probabilities without careful calibration and plenty of meta-uncertainty about how well that calibration extrapolates.

The problem is, most of the value proposition of Jev is that it gives you the probabilities without doing that, which it doesn't.

[0] https://kantahayashiai.github.io/posts/jev-does-not-play-dic...

jezzamon • today at 9:05 PM

Jev should ideally respond in a non-random way, it should just list out the probabilities.

I suppose trying to interface with the model like this is like asking an LLM how many times the letter E appears in a word - just not the correct way to ask that given its model