> the problem is newer models are never trained from scratch
Training a new base model from scratch happens every so often. Closed labs do not publish which models are new base models but as a rule of thumb major release numbers are an indication (with some exceptions).
If the training data is the same, the training algorithms are the same, the RLHF is the same, and the rest of the process is the same, then it's not really from scratch, or not from scratch in a way that results in an 'out of family' model. I doubt any company would take that risk. You always build on and use what works and go from there.
This is true, but Google's models have now had a consistent history of lower psychological* coherence / consistency. See, eg https://arxiv.org/abs/2603.10011 (Gemma Needs Help), or search for recent "Gemini shame loops", where gemini flash models stop producing output other than SHAME SHAME SHAME...
* - as in, Skinner psychology. The set of observable behaviors. Not speaking directly here to anything like an inner life of models.