How does showing the suggested answers to the user make the conversation better for model training?
They could take any conversation without suggested answers, truncate it to just before a user message, have the model predict suggested answers and then train it on the difference between predicted and actual answers, right?
The user actual answer vs the user actual answer after seeing the suggested answer are different points of data
Agreed, they already have the HF -- delta versus model prediction can be calculated at any time.
If anything, showing the suggestion introduces unwanted bias.
RL training, the second phase of LLM training, is based on "I did X, was that good/bad?" and that 1 bit of information is the training data.
So you give the user a suggestion, and the user accepts -> good
You give the user a suggestion, and the user refuses and types something else -> bad (plus some supervisory training data)
The main performance enhancer in LLMs is getting high quality training data. So, first, any extra training data will help. Second this is training data that's directly relevant to their product, and thus higher quality than many other sources.
I'd believe any model provider is mining the shit out of every last customer interaction they can get, not just this.
You can press Tab+Enter to accept it.
The idea is that a thread can have many reasonable follow-ups that the user would've accepted, so it is wrong to punish the model for predicting a follow up that is different from the user message, as that prediction could've been accepted by the user if it was given.