They're estimating a probability distribution over the next token, from which a sample is taken. Close enough.
It's not an estimation of something. It's a policy.
It's not an estimation of something. It's a policy.